DATA CUT 2026-07-27 116 活跃模型 680 A 类案例

专题

Coding Agent 模型指南

从 134 个完整模型中抽取有案例或高分推理/agent 候选;完整集合请看模型索引和对比页。

GLM-4.7

Z AI / GLM

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:13 通用/待核验
已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:12 通用/待核验

Qwen3 Max

Qwen / Alibaba

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:11 通用/待核验

GLM-5.1

Z AI / GLM

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:10 通用/待核验

Solar Pro 3

Upstage / Solar

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:10 通用/待核验

Grok Beta

xAI / Grok

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:10 通用/待核验

Gemini 1.5 Pro (May)

Google / Gemini

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:10 通用/待核验

MiniMax M1 80k

MiniMax

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:9 通用/待核验

Grok 2

xAI / Grok

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:9 通用/待核验

Solar Mini

Upstage / Solar

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:9 通用/待核验

Seedance 1.0

ByteDance Seed

已有真实案例

来自厂商页骨架的补充模型;缺失字段以“暂无数据 / 官方未披露”展示。

A 类案例:9 通用/待核验

Grok 4.3 (high)

xAI / Grok

已有真实案例

来自公开模型资料库,并已补厂商/官方/模型家族证据;当前站内暂无可核验 A 类案例。

A 类案例:8 通用/待核验

可核验案例

116 个模型已有 A 类案例支撑

Braffolk 使用 Claude Fable 5 处理浏览器 3D 世界构建

Braffolk · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:GitHub 最后复核:2026-06-25 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Braffolk 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建浏览器中运行的 4x4km 程序化开放世界,包括地形、植被、光照、云、QA 和性能验证

公开产物:公开 GitHub repo、公开 demo、README 说明项目约 99% 由 Fable 5 构建

模型作用:README 称模型完成架构、引擎、世界系统、验证工具、调试和工作记录

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:模型贡献比例来自作者自述;需要保存 README、HEAD SHA、demo 截图和 STATUS.md 快照

原始记录:LAAS: 4x4km procedural WebGPU open world

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

StarKnightt 使用 Claude Fable 5 处理浏览器 3D 世界构建

StarKnightt · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:Reddit; GitHub; Vercel 最后复核:2026-06-25 证据快照:已快照 · 3 / 3 个证据目标 归档说明:已完成快照:3 / 3 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据2 个公开产物复核通过代码仓库证据

StarKnightt 公开的3D 与 Web 交互案例,来源为 公开代码库、社区原帖、在线 Demo,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建可玩的第一人称 Backrooms survival horror 浏览器游戏

公开产物:公开 Reddit 原帖、GitHub repo、Vercel demo;README 描述玩法、程序化资产、声音、移动端和测试

模型作用:Reddit 标题和 repo 描述均指向 Claude Fable 5;具体贡献比例仍来自作者自述

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:需要保存 Reddit 原帖和 demo 运行截图;模型参与程度无第三方独立复核

原始记录:Backrooms Escape browser horror game

已有真实案例 3D 与 Web 交互公开代码库、社区原帖、在线 DemoA 类可核验real_case auto_approved

hiImMate 使用 Claude Fable 5 处理可玩交互原型构建

hiImMate · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:Reddit; Vercel 最后复核:2026-06-25 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

hiImMate 公开的游戏与交互原型案例,来源为 社区原帖、在线 Demo,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建免费 survivor browser game,并带有作者描述的多人联机能力

公开产物:Reddit 原帖提供 playable Vercel demo;作者称 “made entirely by Fable 5”

模型作用:作者称自己没有查看代码,项目完全由 Fable 5 制作

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:未找到 repo;artifact 为 live demo,必须保存截图/录屏;多人联机能力需人工复测

原始记录:Cube Survivor multiplayer survivor game

已有真实案例 游戏与交互原型社区原帖、在线 DemoA 类可核验real_case auto_approved

next-choken / levy-street 使用 Claude Fable 5 处理可玩交互原型构建

next-choken / levy-street · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:Reddit; GitHub; product_site 最后复核:2026-06-25 证据快照:已快照 · 3 / 3 个证据目标 归档说明:已完成快照:3 / 3 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据2 个公开产物复核通过代码仓库证据

next-choken / levy-street 公开的游戏与交互原型案例,来源为 公开代码库、社区原帖、产品页面,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:用 Fable 5 构建可玩的 browser MMORPG,包括任务、职业、战斗、多人、服务器和开源代码

公开产物:Reddit 原帖、公开网站、公开 GitHub repo;README 说明 playable browser MMO、self-host 和 AI agent training

模型作用:Reddit 原帖明确 “Used Fable to build a full blown MMORPG”

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目规模较大,需快照 repo HEAD、README 和网站首屏;资产来源需在正式展示前复核

原始记录:World of ClaudeCraft browser MMORPG

已有真实案例 游戏与交互原型公开代码库、社区原帖、产品页面A 类可核验real_case auto_approved 进入模型卡精选

PlayfulInterview984 使用 Claude Fable 5 处理可玩交互原型构建

PlayfulInterview984 · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:Reddit; project_site; YouTube 最后复核:2026-06-25 证据快照:待快照 · 0 / 3 个证据目标 归档说明:已列入归档队列,等待抓取 3 个证据目标。
证据可信度 100/100

A+ 完整链路 · 社区公开记录

原始证据2 个公开产物复核通过社区公开记录

PlayfulInterview984 公开的游戏与交互原型案例,来源为 社区原帖、视频证据,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:对 1989 年 DOS 游戏可执行文件进行解包、反汇编、函数映射、算法复现和验证

公开产物:Reddit 原帖提供视频和 playable tech demo;作者称 Fable 5 overnight 解码 602 个函数并复现 terrain generator

模型作用:作者描述 Fable 5 通过 parallel agents、evidence ledger 和 bit-for-bit 验证完成关键解码

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:涉及旧游戏 IP 和逆向工程,展示时必须保留合法性/授权风险备注;视频需记录关键时间点

原始记录:Midwinter 1989 DOS executable decoding and remaster workflow

已有真实案例 游戏与交互原型社区原帖、视频证据A 类可核验real_case auto_approved 进入模型卡精选

huiung 使用 Claude Fable 5 处理可玩交互原型构建

huiung · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:GitHub; product_site; video 最后复核:2026-06-25 证据快照:已快照 · 3 / 3 个证据目标 归档说明:已完成快照:3 / 3 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据2 个公开产物复核通过代码仓库证据

huiung 公开的游戏与交互原型案例,来源为 公开代码库、产品页面,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建浏览器中的空间模拟 / MMO 原型,包含飞行、采矿、交易、战斗、聊天和多人

公开产物:公开 repo、play now 网站、showcase video;README 声称 Built with Claude Fable 5

模型作用:repo 描述和 README 表明项目由 Claude Fable 5 构建

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:使用 “Star Citizen” 作比较对象,需注意商标/IP 风险;模型贡献仍为作者/README 自述

原始记录:Claude Citizen browser space MMO prototype

已有真实案例 游戏与交互原型公开代码库、产品页面A 类可核验real_case auto_approved

Randroids-Dojo 使用 Claude Fable 5 处理可玩交互原型构建

Randroids-Dojo · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:GitHub; Vercel 最后复核:2026-06-25 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Randroids-Dojo 公开的游戏与交互原型案例,来源为 公开代码库、在线 Demo,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建 web-native pinball game

公开产物:公开 GitHub repo、公开 Vercel demo;README 声称 Built with Claude Fable 5 while on stream

模型作用:repo description 和 README 直接声明 Fable 5 参与构建

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:repo star/activity 很低,需保存 demo 截图并尽量补直播原始链接

原始记录:VibePinball web-native pinball game

已有真实案例 游戏与交互原型公开代码库、在线 DemoA 类可核验real_case auto_approved

Pluventi / kijai/ComfyUI-KJNodes 使用 Claude Fable 5 处理多模态内容处理

Pluventi / kijai/ComfyUI-KJNodes · Claude Fable 5

A
厂商:Anthropic / Claude 模型:Claude Fable 5 来源平台:GitHub PR 最后复核:2026-06-25 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Pluventi / kijai/ComfyUI-KJNodes 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-25。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:为 ComfyUI KJNodes 添加 Ideogram 4 Prompt Builder V2,包括 freehand drawing、canvas tools、layer 管理和节点文件

公开产物:公开 PR、commit list、作者评论和后续引用;评论中写明 Built with Claude Fable 5

模型作用:作者评论称自己不会写代码、AI 写了 100% 代码,后续评论说明 Built with Claude Fable 5

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:PR 质量作者自己也提示需开发者清理;进入精选时适合展示 “AI coding with human testing”,不是质量保证案例

原始记录:Ideogram 4 Prompt Builder KJ V2 PR for ComfyUI-KJNodes

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Khan Academy 使用 GPT-4 处理智能体流程编排

Khan Academy · GPT-4

A
厂商:OpenAI 模型:GPT-4 来源平台:OpenAI story; Khan Academy product 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Khan Academy 公开的智能体工作流案例,来源为 OpenAI story; Khan Academy product,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Khan Academy used GPT-4 to power Khanmigo, an AI tutor for students and assistant for teachers.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public OpenAI story and live Khanmigo product page describe the deployed learning assistant.

模型作用:GPT-4 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI states Khanmigo uses GPT-4; Khan Academy product page describes the tutor and teacher assistant workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Education outcomes and safety claims still need independent outcome studies; model version may have evolved after launch.

原始记录:Khanmigo AI tutor and teaching assistant

已有真实案例 智能体工作流OpenAI story; Khan Academy productA 类可核验real_case auto_approved 进入模型卡精选

Duolingo 使用 GPT-4 处理真实任务执行

Duolingo · GPT-4

A
厂商:OpenAI 模型:GPT-4 来源平台:OpenAI story; Duolingo blog 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Duolingo 公开的真实任务执行案例,来源为 博客记录,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Duolingo used GPT-4 in Duolingo Max for Roleplay and Explain My Answer learning features.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public OpenAI story and Duolingo launch post describe the product tier and user-facing AI features.

模型作用:GPT-4 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI and Duolingo both identify GPT-4 as the model powering the two Max features.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Subscription packaging and model routing can change; keep current product availability separate from original GPT-4 launch evidence.

原始记录:Duolingo Max Roleplay and Explain My Answer

已有真实案例 真实任务执行博客记录A 类可核验real_case auto_approved 进入模型卡精选

Be My Eyes 使用 GPT-4 处理知识检索和问答

Be My Eyes · GPT-4

A
厂商:OpenAI 模型:GPT-4 来源平台:OpenAI story; Be My Eyes news 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Be My Eyes 公开的知识库与检索问答案例,来源为 OpenAI story; Be My Eyes news,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Be My Eyes built Be My AI / Virtual Volunteer to answer image questions for blind and low-vision users.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public OpenAI story and Be My Eyes announcement describe the GPT-4-powered app feature.

模型作用:GPT-4 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Both sources identify GPT-4 visual capability as the model layer behind the assistant.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Accessibility-critical outputs require careful accuracy and safety review; product model routing may change over time.

原始记录:Be My AI visual accessibility assistant

已有真实案例 知识库与检索问答OpenAI story; Be My Eyes newsA 类可核验real_case auto_approved 进入模型卡精选

Morgan Stanley 使用 GPT-4 处理知识检索和问答

Morgan Stanley · GPT-4

A
厂商:OpenAI 模型:GPT-4 来源平台:OpenAI story; Morgan Stanley press release 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Morgan Stanley 公开的知识库与检索问答案例,来源为 OpenAI story; Morgan Stanley press release,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Morgan Stanley used GPT-4 to help financial advisors retrieve and summarize internal knowledge.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:OpenAI story and Morgan Stanley release describe the advisor-facing internal knowledge system.

模型作用:GPT-4 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Morgan Stanley states its wealth management solution uses GPT-4 on internal Morgan Stanley content.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Internal enterprise artifact is not publicly usable; evidence relies on official company description rather than an inspectable demo.

原始记录:Morgan Stanley Wealth Management knowledge assistant

已有真实案例 知识库与检索问答OpenAI story; Morgan Stanley press releaseA 类可核验real_case auto_approved 进入模型卡精选

Stripe 使用 GPT-4 处理智能体流程编排

Stripe · GPT-4

A
厂商:OpenAI 模型:GPT-4 来源平台:OpenAI story; Stripe newsroom 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Stripe 公开的智能体工作流案例,来源为 OpenAI story; Stripe newsroom,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Stripe used GPT-4 to improve developer support, business understanding and fraud detection workflows.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:OpenAI story and Stripe newsroom post describe GPT-4-powered prototypes and platform workflows.

模型作用:GPT-4 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI and Stripe identify GPT-4 as the model used in the Stripe product and operations exploration.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Some workflows are internal or prototype-stage; keep separated from user-visible product claims.

原始记录:Stripe GPT-4 developer support and fraud workflows

已有真实案例 智能体工作流OpenAI story; Stripe newsroomA 类可核验real_case auto_approved

Be My Eyes 使用 GPT-4o (May) 处理多模态内容处理

Be My Eyes · GPT-4o (May)

A
厂商:OpenAI 模型:GPT-4o (May) 来源平台:OpenAI launch; YouTube; product site 最后复核:2026-06-26 证据快照:待快照 · 0 / 3 个证据目标 归档说明:已列入归档队列,等待抓取 3 个证据目标。
证据可信度 100/100

A+ 完整链路 · 官方/客户故事

原始证据2 个公开产物复核通过官方/客户故事

Be My Eyes 公开的多模态生成与理解案例,来源为 视频证据,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Be My Eyes demonstrated GPT-4o assisting a blind user with real-time visual understanding in a live accessibility workflow.

公开产物:公开材料提供视频记录、页面说明或可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:OpenAI launch page and Be My Eyes demo video provide the concrete model, user scenario and visible artifact.

模型作用:GPT-4o (May) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI presented the Be My Eyes accessibility workflow as a GPT-4o multimodal capability demo.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a launch/demo artifact rather than a full production case study; verify current product routing before claiming ongoing GPT-4o use.

原始记录:Be My Eyes live accessibility demo with GPT-4o

已有真实案例 多模态生成与理解视频证据A 类可核验real_case auto_approved 进入模型卡精选

Replit 使用 Claude 4 Opus 处理软件工程任务执行

Replit · Claude 4 Opus

A
厂商:Anthropic / Claude 模型:Claude 4 Opus 来源平台:official_web 最后复核:2026-06-27T08:47:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Replit 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:47:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Use Claude Opus 4 in Replit's AI coding environment to make precise, complex changes spanning multiple files.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic's launch evidence reports improved precision and dramatic advancements for Replit on complex multi-file changes; Replit's public AI product page is reachable as the artifact.

模型作用:Claude 4 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4 supplies stronger coding precision and long-horizon multi-file reasoning for Replit's software-building agent workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an Anthropic launch-page customer statement; artifact is Replit's public AI product page.

原始记录:Replit applies Claude Opus 4 to complex multi-file code changes

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

SOMIN.ai 使用 Gemini 2.5 Pro 处理研究分析和报告生成

SOMIN.ai · Gemini 2.5 Pro

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro 来源平台:Google Cloud customer story; product site 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

SOMIN.ai 公开的研究与报告生成案例,来源为 Google Cloud customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SOMIN.ai used Gemini 2.5 Pro to analyze public marketing data and audience research for campaign planning.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Google Cloud customer story gives model, user, workflow and quantified product outcome.

模型作用:Gemini 2.5 Pro 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:The case states Gemini 2.5 Pro completes analysis and idea generation for SOMIN workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Business impact claims come from vendor/customer story; keep as official case evidence, not independent benchmark.

原始记录:SOMIN campaign planning and creative analysis

已有真实案例 研究与报告生成Google Cloud customer story; product siteA 类可核验real_case auto_approved 进入模型卡精选

Sofya 使用 Llama 3.1 405B 处理智能体流程编排

Sofya · Llama 3.1 405B

A
厂商:Meta / Llama 模型:Llama 3.1 405B 来源平台:Llama case study 最后复核:2026-06-26 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Sofya 公开的智能体工作流案例,来源为 Llama case study,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Sofya used Llama 3.1 405B to generate high-quality synthetic training data for its solution.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Llama case study provides model, organization, task and implementation narrative.

模型作用:Llama 3.1 405B 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:The case identifies Llama 3.1 405B as part of the synthetic data generation and distillation workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The public artifact is a case-study page rather than a runnable demo; treat as official customer-case evidence.

原始记录:Sofya synthetic training data generation

已有真实案例 智能体工作流Llama case studyA 类可核验real_case auto_approved 进入模型卡精选

Perplexity 使用 Claude 3.5 Sonnet 处理知识检索和问答

Perplexity · Claude 3.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.5 Sonnet 来源平台:Anthropic customer story; product site 最后复核:2026-06-26 证据快照:需关注 · 0 / 2 个证据目标 归档说明:当前快照需要关注,请检查证据指纹、URL 或 HTTP 状态。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Perplexity 公开的知识库与检索问答案例,来源为 Anthropic customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Perplexity used Claude 3.5 Sonnet for complex, context-sensitive paid-user answer workflows.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story describes the Claude-powered search product and model-role split.

模型作用:Claude 3.5 Sonnet 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Anthropic states Perplexity uses Claude 3.5 Sonnet in its paid offering for advanced reasoning and top performance.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Perplexity offers multiple models; this case proves availability/use, not exclusive model routing.

原始记录:Perplexity answer engine Claude model routing

已有真实案例 知识库与检索问答Anthropic customer story; product siteA 类可核验real_case auto_approved 进入模型卡精选

Triple Whale 使用 Claude 3.7 Sonnet 处理研究分析和报告生成

Triple Whale · Claude 3.7 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.7 Sonnet 来源平台:Anthropic customer story; product site 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Triple Whale 公开的研究与报告生成案例,来源为 Anthropic customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Triple Whale selected Claude 3.7 Sonnet as the default agent model for complex ecommerce data analysis.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story and Triple Whale product page provide model, user, task and product evidence.

模型作用:Claude 3.7 Sonnet 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Anthropic states Claude 3.7 Sonnet was selected for context retention and accuracy on complex numerical datasets.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Business impact claims come from vendor/customer story; retain as official case evidence, not independent benchmark.

原始记录:Triple Whale Moby ecommerce intelligence agent

已有真实案例 研究与报告生成Anthropic customer story; product siteA 类可核验real_case auto_approved 进入模型卡精选

Grafana Labs 使用 Claude 3.7 Sonnet 处理研究分析和报告生成

Grafana Labs · Claude 3.7 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.7 Sonnet 来源平台:Anthropic customer story; Grafana product 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Grafana Labs 公开的研究与报告生成案例,来源为 Anthropic customer story; Grafana product,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Grafana used Claude Sonnet 3.7 in Grafana Assistant for technically complex observability tasks.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story describes the assistant, model family and division of tasks.

模型作用:Claude 3.7 Sonnet 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:The case quotes Grafana using Claude Sonnet 3.7 for technically complex tasks while Haiku handles simpler summaries.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Grafana notes a transition toward Sonnet 4; keep this as 3.7-era model evidence.

原始记录:Grafana Assistant observability workflows

已有真实案例 研究与报告生成Anthropic customer story; Grafana productA 类可核验real_case auto_approved 进入模型卡精选

Semgrep 使用 Claude 3.7 Sonnet 处理研究分析和报告生成

Semgrep · Claude 3.7 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.7 Sonnet 来源平台:Anthropic customer story; Semgrep product 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Semgrep 公开的研究与报告生成案例,来源为 Anthropic customer story; Semgrep product,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Semgrep used Claude 3.7 Sonnet for deeper contextual understanding in code vulnerability analysis.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story and Semgrep product page identify the security assistant workflow.

模型作用:Claude 3.7 Sonnet 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:The case says Claude 3.7 Sonnet recognized a vulnerability in generated code and improved key evals.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Security claims need periodic review against current product model routing and evaluation methodology.

原始记录:Semgrep code security analysis

已有真实案例 研究与报告生成Anthropic customer story; Semgrep productA 类可核验real_case auto_approved

Vanta 使用 Claude 3.7 Sonnet 处理软件工程任务执行

Vanta · Claude 3.7 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.7 Sonnet 来源平台:Anthropic customer story; product site 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Vanta 公开的代码代理与软件工程案例,来源为 Anthropic customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Vanta used Claude 3.7 Sonnet through Cursor for code generation and remediation workflows.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story provides model, tool chain, organization and engineering task evidence.

模型作用:Claude 3.7 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The case quotes Vanta on Claude 3.7 Sonnet code quality building trust among engineers.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Artifact is an enterprise workflow, not a public repo; evidence is official customer-story level.

原始记录:Vanta engineering remediation with Cursor

已有真实案例 代码代理与软件工程Anthropic customer story; product siteA 类可核验real_case auto_approved

Steno 使用 Claude 3 Opus 处理研究分析和报告生成

Steno · Claude 3 Opus

A
厂商:Anthropic / Claude 模型:Claude 3 Opus 来源平台:Anthropic customer story; product site 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Steno 公开的研究与报告生成案例,来源为 Anthropic customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Steno used Claude 3 Opus to help attorneys find relevant information across deposition transcripts.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story identifies Claude 3 Opus, the legal workflow and the Steno product context.

模型作用:Claude 3 Opus 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:The case states Claude helps lawyers quickly find relevant information across vast transcripts.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Legal workflows require accuracy review and human oversight; product internals are not publicly inspectable.

原始记录:Steno deposition preparation assistant

已有真实案例 研究与报告生成Anthropic customer story; product siteA 类可核验real_case auto_approved 进入模型卡精选

Sourcegraph 使用 Claude 3 Opus 处理知识检索和问答

Sourcegraph · Claude 3 Opus

A
厂商:Anthropic / Claude 模型:Claude 3 Opus 来源平台:Anthropic customer story; Sourcegraph product 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Sourcegraph 公开的知识库与检索问答案例,来源为 Anthropic customer story; Sourcegraph product,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Sourcegraph used Claude 3 Opus to improve Cody's understanding of large code contexts.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story and Sourcegraph Cody product page identify the coding assistant artifact.

模型作用:Claude 3 Opus 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:The case attributes improved large-context code understanding and recall to Claude 3 Opus.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Cody supports multiple model providers; this case should be read as one Claude-backed configuration.

原始记录:Sourcegraph Cody large-codebase context

已有真实案例 知识库与检索问答Anthropic customer story; Sourcegraph productA 类可核验real_case auto_approved 进入模型卡精选

Clay 使用 Claude 3 Haiku 处理智能体流程编排

Clay · Claude 3 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3 Haiku 来源平台:Anthropic customer story; product site 最后复核:2026-06-26 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Clay 公开的智能体工作流案例,来源为 Anthropic customer story; product site,复核于 2026-06-26。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Clay used Claude 3 Haiku for lead identification, enrichment and personalized sales messaging.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic customer story connects the Claude Haiku family to live Clay growth workflows.

模型作用:Claude 3 Haiku 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:The story identifies Claude 3 Haiku as a top choice model for cost-effective outreach tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Clay may route across models as Anthropic updates the Claude family; keep this as dated Claude 3 Haiku evidence.

原始记录:Clay sales enrichment and messaging

已有真实案例 智能体工作流Anthropic customer story; product siteA 类可核验real_case auto_approved 进入模型卡精选

Replit 使用 Claude 3.5 Sonnet 处理代码审查和测试生成

Replit · Claude 3.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.5 Sonnet 来源平台:official_customer_story 最后复核:2026-06-26T20:20:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Replit 公开的代码审查与测试案例,来源为 官方页面,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Replit, an agentic software creation platform, uses Claude to help anyone build and deploy software applications without coding experience. Michele Catasta, President and Head of AI, stated: 'We made the choice back whe…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Replit reached 50 million+ users building applications without coding skills. Annual recurring revenue grew from $240M. Teams from Zillow to Duolingo deploy Agent 4 to ship internal tools. One user generated 36,000+ lin…

模型作用:Claude 3.5 Sonnet 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Sonnet was the foundational model Replit chose to build their AI agent upon in 2024. The model's reasoning depth, vision capabilities for self-testing, and long-running task consistency enabled Agent 4 to man…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Customer story is authored by Anthropic, with quotes from Replit's President. The platform now uses newer Sonnet/Opus versions for Agent 4, but the foundational choice was made with Sonnet 3.5.

原始记录:Replit selects Claude 3.5 Sonnet as foundation for Agent 4, reaching 50M+ users without coding experience

已有真实案例 代码审查与测试官方页面A 类可核验real_case auto_approved 进入模型卡精选

Lovable (Anton Osika, CEO) 使用 Claude 3.5 Sonnet 处理软件工程任务执行

Lovable (Anton Osika, CEO) · Claude 3.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.5 Sonnet 来源平台:official_customer_story 最后复核:2026-06-26T20:20:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Lovable (Anton Osika, CEO) 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Lovable turns plain-language descriptions of an app into production-grade software through conversational refinement. Alexandre Pesant, product lead, stated: 'Claude Sonnet 3.5 was the first model that made agents work.…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Within a year of launch, Lovable reached $200M in annualized recurring revenue. Over 50 million projects have been built on the platform (200,000+ per day), drawing more than 600 million visits per month. Non-technical …

模型作用:Claude 3.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Sonnet was specifically identified as 'the first model that made agents work' by Lovable's product lead. It enabled the transition from simple instruction-following to autonomous agentic software generation w…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Customer story authored by Anthropic. Quote attribution is to Lovable's product lead Alexandre Pesant. Later Opus 4.5 is described as the next step change, but 3.5 Sonnet was the initial breakthrough.

原始记录:Lovable credits Claude 3.5 Sonnet as first model enabling agents, powering 50M+ app builds and $200M ARR

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Dorian (Doriandarko, GitHub) 使用 Claude 3.5 Sonnet 处理软件工程任务执行

Dorian (Doriandarko, GitHub) · Claude 3.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.5 Sonnet 来源平台:github 最后复核:2026-06-26T20:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Dorian (Doriandarko, GitHub) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Claude Engineer is an interactive command-line interface (CLI) that leverages Claude 3.5 Sonnet to assist with software development tasks. The framework enables Claude to generate and manage its own tools, continuously …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project reached 11,196 GitHub stars and 1,167 forks, becoming one of the most popular open-source tools built on Claude 3.5 Sonnet. It enables developers to use Claude for autonomous code generation, file manipulati…

模型作用:Claude 3.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Sonnet is the core and only model powering the tool. Its instruction following, code generation, and tool-use capabilities enable the CLI to autonomously create and manage development tools, making it the fou…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source project by individual developer. Stars/forks indicate community adoption but not necessarily enterprise production deployment. The tool was specifically designed for Claude 3.5 Sonnet.

原始记录:Claude Engineer: 11K-star open-source CLI leveraging Claude 3.5 Sonnet for autonomous software development assistance

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Murat Can Koylan 使用 Claude 3.5 Sonnet 处理软件工程任务执行

Murat Can Koylan · Claude 3.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.5 Sonnet 来源平台:github 最后复核:2026-06-26T20:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Murat Can Koylan 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AI-Investigator is an automated Python framework designed to analyze website content and generate structured reports using Claude 3.5 Sonnet API (claude-3-5-sonnet-20241022) and Firecrawl. The system supports two main m…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A functional Python framework that crawls websites, extracts content, and uses Claude 3.5 Sonnet to analyze and generate structured reports on enterprise AI case studies. The system achieved 722 GitHub stars and demonst…

模型作用:Claude 3.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Sonnet serves as the core intelligence engine, handling content analysis, entity extraction, and structured report generation. The model's reasoning capabilities enable it to identify relevant case studies, e…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Individual developer project. Lower star count (722) compared to major tools. Uses a specific dated model version (claude-3-5-sonnet-20241022). Architecture is well-documented with clear Claude integration.

原始记录:AI-Investigator: Automated enterprise content analysis framework built on Claude 3.5 Sonnet and Firecrawl

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Simon Willison (独立开发者,Datasette 作者) 使用 Claude 3 Opus 处理真实任务执行

Simon Willison (独立开发者,Datasette 作者) · Claude 3 Opus

A
厂商:Anthropic / Claude 模型:Claude 3 Opus 来源平台:GitHub 最后复核:2026-06-26T20:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Simon Willison (独立开发者,Datasette 作者) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建一个 Python CLI 工具 files-to-prompt,将目录中的多个文件拼接成单个 prompt 供 LLM(如 Claude 和 GPT-4)使用。整个开发过程(初始骨架生成、功能迭代、自我迭代改进)几乎完全由 Claude 3 Opus 完成。

公开产物:成功发布 files-to-prompt 工具(GitHub 2753+ stars),可将任意目录文件序列化为适合长上下文模型的 prompt,与 Simon Willison 的 llm CLI 工具配合使用。工具支持多种文件格式和递归目录遍历。

模型作用:Claude 3 Opus 是整个项目的主要编码工具:从 cookiecutter 模板生成初始骨架、编写核心逻辑、处理 edge cases、编写测试,以及在工具可用后用于自我迭代改进。Simon Willison 明确表示 'I wrote files-to-prompt almost entirely using Claude 3 Opus'。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人开发者项目,非企业级生产部署;但项目影响力大(Simon Willison 是知名开发者,GitHub 2753 stars)。

原始记录:Simon Willison 构建 files-to-prompt CLI 工具,全程使用 Claude 3 Opus 编码

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Matt Shumer (HyperWrite AI 创始人) 使用 Claude 3 Opus 处理研究分析和报告生成

Matt Shumer (HyperWrite AI 创始人) · Claude 3 Opus

A
厂商:Anthropic / Claude 模型:Claude 3 Opus 来源平台:GitHub 最后复核:2026-06-26T20:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Matt Shumer (HyperWrite AI 创始人) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建 Claude-Investor,一个使用 Claude 3 Opus 和 Haiku 模型的实验性投资分析 Agent。能够自动获取指定行业的公司列表、历史价格数据、财务报表、新闻文章,执行情感分析,检索分析师评级,并生成综合比较分析报告和投资建议。

公开产物:成功发布 Claude-Investor 项目(GitHub 2258+ stars),实现了完整的投资分析流水线:从行业筛选、数据获取、情感分析到生成包含价格目标的综合投资报告。项目以 Jupyter Notebook 形式提供,可直接在 Google Colab 中运行。

模型作用:Claude 3 Opus 作为核心推理引擎,负责多步复杂分析任务:跨多个数据源的信息整合、财务报表解读、情感分析、行业比较分析和投资建议生成。Haiku 用于成本敏感的辅助任务。Opus 的长上下文能力使其能同时处理大量财务数据和新闻文章。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:实验性项目,免责声明明确指出不构成投资建议;Matt Shumer 是知名 AI 创业者(HyperWrite AI),项目有一定影响力。

原始记录:Matt Shumer 使用 Claude 3 Opus 构建 Claude-Investor 智能投资分析 Agent

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Seth Hobson (独立开发者) 使用 Claude 3 Opus 处理研究分析和报告生成

Seth Hobson (独立开发者) · Claude 3 Opus

A
厂商:Anthropic / Claude 模型:Claude 3 Opus 来源平台:GitHub 最后复核:2026-06-26T20:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Seth Hobson (独立开发者) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建一个金融对话应用,使用 LangChain、LangGraph、OpenBB 和 Claude 3 Opus 实现智能股票分析。功能包括:通过 OpenBB 获取金融数据、生成技术分析摘要、股票价格历史和定量统计、相对强度计算、新闻情感分析、基于 FinViz 的宇宙扫描、基于技术指标的风险管理,以及多 Agent 工作流。

公开产物:成功发布 financial-chat 应用(GitHub 235 stars),支持 Streamlit UI 和 FastAPI 两种交互方式。项目包含完整的多 Agent 工作流(通过 LangGraph 实现),可部署到 AWS。作者撰写了一系列详细的博客文章记录开发过程。

模型作用:Claude 3 Opus 作为核心 LLM 引擎,驱动整个 Agent 系统的推理和决策:理解用户自然语言查询、调用 OpenBB 工具获取金融数据、生成技术分析报告、执行多步推理完成复杂分析任务。Opus 的长上下文和工具调用能力是实现复杂 Agent 工作流的关键。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人项目,非企业级生产部署;但有完整的博客系列和部署文档,技术实现质量较高。

原始记录:Seth Hobson 构建基于 LangChain + Claude 3 Opus 的智能股票分析对话工具

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Burak Bekci (rakyll, Google 工程师,Go … 使用 Claude 3 Opus 处理真实任务执行

Burak Bekci (rakyll, Google 工程师,Go pprof contributor) · Claude 3 Opus

A
厂商:Anthropic / Claude 模型:Claude 3 Opus 来源平台:GitHub 最后复核:2026-06-26T20:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Burak Bekci (rakyll, Google 工程师,Go pprof contributor) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T20:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建 kubehelp,一个命令行工具,用户输入自然语言描述(如 'deploy helloworld to the test namespace and expose it at 8080'),Claude 3 Opus 将其转换为对应的 kubectl 命令并可选择直接执行。支持创建部署、暴露服务、查询日志、删除资源等常见 Kubernetes 操作。

公开产物:成功发布 kubehelp 工具(GitHub 63 stars),可将自然语言描述准确转换为 kubectl 命令,支持创建 namespace/deployment/service、查看日志、查询集群信息、删除资源等操作。工具使用 Go 编写,通过 Anthropic API 调用 Claude 3 Opus。

模型作用:Claude 3 Opus 作为核心翻译引擎,负责理解用户的自然语言 Kubernetes 操作意图,并准确生成对应的 kubectl 命令。Opus 的代码生成能力和对 Kubernetes 概念的深入理解使其能够处理复杂的多步操作(如先创建 namespace 再部署应用再暴露服务)。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:开发者工具,面向 Kubernetes 用户;作者是知名 Go 开发者(rakyll),项目技术质量高但用户规模较小。

原始记录:Burak Bekci 使用 Claude 3 Opus 构建 kubehelp — 自然语言生成 kubectl 命令

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AWS Samples (AWS Hackathon France 2… 使用 Claude 3 Haiku 处理研究分析和报告生成

AWS Samples (AWS Hackathon France 2024 winner) · Claude 3 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3 Haiku 来源平台:github 最后复核:2026-06-26T20:15:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AWS Samples (AWS Hackathon France 2024 winner) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T20:15:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Personalized nutritional analysis of grocery products via barcode scanning and food photo recognition; generate ingredient explanations, quantitative allergen detection, and personalized recipe suggestions based on user…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production serverless app deployed on AWS (Lambda + DynamoDB + S3 + CloudFront + Bedrock); winner of AWS Hackathon France 2024; showcased at multiple AWS Summit events; processes product barcodes and food photos with pe…

模型作用:Claude 3 Haiku 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3 Haiku was explicitly chosen as 'a good ratio between latency and quality' for real-time food ingredient analysis, allergen detection, and recipe generation. Handles multi-shot prompt engineering for generating …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:AWS sample/demo project built for hackathon and summit demos, not a commercial product with paying customers. However, it is a real deployed application with production-grade serverless architecture and demonstrates Cla…

原始记录:AWS Food Analyzer: Serverless Nutrition App Uses Claude 3 Haiku for Real-Time Ingredient Analysis and Recipe Generation

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

VizuaraAI 使用 Gemini 2.5 Pro 处理软件工程任务执行

VizuaraAI · Gemini 2.5 Pro

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro 来源平台:github 最后复核:2026-06-27T09:34:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

VizuaraAI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:34:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Paper2Notebook lets users upload a research-paper PDF or provide an arXiv URL, then generates a runnable Jupyter notebook with paper analysis, methodology, PyTorch implementation, experiments, validation, and Colab-read…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains a FastAPI backend and Next.js frontend for producing downloadable executable notebooks; its README describes PDF-to-notebook conversion, LaTeX extraction, PyTorch implementation, dependenc…

模型作用:Gemini 2.5 Pro 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro is the default model in the backend (DEFAULT_MODEL = "gemini-2.5-pro") and is called for paper analysis, implementation design, code-cell generation, JSON repair, and validation of the generated notebook …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository README and backend source code; deployment volume and production usage metrics are not disclosed.

原始记录:VizuaraAI Paper2Notebook uses Gemini 2.5 Pro to convert research PDFs into executable Jupyter notebooks

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hylarucoder 使用 Gemini 2.5 Pro 处理智能体流程编排

hylarucoder · Gemini 2.5 Pro

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro 来源平台:github 最后复核:2026-06-26T12:30:20Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hylarucoder 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T12:30:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Use Gemini 2.5 Pro's reasoning capability to detect and rewrite AI-generated Chinese text patterns ('AI味') to produce human-sounding content. The project provides a structured prompt that instructs the model to rewrite articles while preserving meaning, reducing AI detection scores from ~70% to ~17% as measured by 朱雀大模型 detector.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Open-source prompt methodology (1,067 GitHub stars) validated specifically on Gemini 2.5 Pro. Achieves: 1000→2000 word expansion with <22% AI detection increase; 5000-word rewrite reducing AI detection from ~70% to ~17%…

模型作用:Gemini 2.5 Pro is the only model validated for this workflow — the project explicitly states '仅在 Gemini 2.5 Pro 上测试通过' (only tested on Gemini 2.5 Pro) and requires a reasoning model. The model's Chinese language fluency, long-context understanding, and reasoning ability are essential for the nuanced rewriting task.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Prompt-based project, not a deployed product. Effectiveness depends on detection tool used (朱雀大模型). Author notes results may vary.

原始记录:hylarucoder AI味去除: Chinese AI Text Humanization on Gemini 2.5 Pro

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

raizamartin 使用 Gemini 2.5 Pro 处理知识检索和问答

raizamartin · Gemini 2.5 Pro

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro 来源平台:github 最后复核:2026-06-26T12:30:20Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

raizamartin 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T12:30:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an interactive terminal-based AI coding assistant that uses Gemini 2.5 Pro as its primary model for code generation, file operations (view/edit/list/grep), directory management, system commands, linting, and test …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Published PyPI package 'gemini-code' (546 GitHub stars). Full-featured CLI tool with markdown rendering, chat sessions, history management, and automatic tool usage. Includes Notion documentation page. Python codebase w…

模型作用:Gemini 2.5 Pro 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro is the flagship model supported and explicitly named in the project. The tool leverages the model's code generation, tool calling, and instruction following capabilities for autonomous file editing, bash …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Individual developer project. README states 'powered by Gemini 2.5 Pro' but also supports other models. PyPI package was not independently verified during this search.

原始记录:Gemini Code: Terminal AI Coding Assistant Powered by Gemini 2.5 Pro

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Cedrick Chee 使用 Gemini 2.5 Pro (Mar) 处理浏览器 3D 世界构建

Cedrick Chee · Gemini 2.5 Pro (Mar)

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro (Mar) 来源平台:github 最后复核:2026-06-27T07:53:33Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Cedrick Chee 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-27T07:53:33Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕浏览器 3D 世界构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a browser-based 3D multiplayer flying game with arcade-style mechanics and document the AI-assisted development process from March 26-27, 2025.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public repository for Vibe Jet with screenshots, videos, source code and development notes; the README states the project was created using the Gemini 2.5 Pro experimental model via vibe coding.

模型作用:Gemini 2.5 Pro (Mar) 在该案例中承担浏览器 3D 世界构建相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro Experimental provided the coding assistance and initial working foundation for the game, including feature recommendations and code-generation support during development.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence comes from the project author's README, not a third-party audit; it explicitly names Gemini 2.5 Pro experimental and includes the public artifact repository.

原始记录:Cedrick Chee built the Vibe Jet browser 3D multiplayer flying game with Gemini 2.5 Pro Experimental

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Alex Safari (Alex-safari) 使用 Gemini 2.5 Pro 处理研究分析和报告生成

Alex Safari (Alex-safari) · Gemini 2.5 Pro

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro 来源平台:github 最后复核:2026-06-26T12:30:20Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Alex Safari (Alex-safari) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T12:30:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Automated n8n workflow that converts a single product photo into a Hollywood-quality video ad. Gemini 2.5 Pro is used for creative direction — generating cinematic timeline prompts with vibe/format/genre, 12-second shot…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Complete n8n automation workflow (56 GitHub stars) with detailed documentation. Pipeline: product photo upload → GPT-4o image analysis → Gemini 2.5 Pro creative direction → Sora 2 video generation → Google Drive storage…

模型作用:Gemini 2.5 Pro 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro serves as the 'film director' in the pipeline — it takes structured product analysis data and generates professional-grade cinematic prompts for video generation. The model's creative writing, structured …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Workflow depends on multiple paid services (OpenRouter, Sora 2 via Kie.ai, Google Drive). Gemini 2.5 Pro is one component in a multi-model pipeline, not the sole model.

原始记录:Hollywood-Quality UGC Ad Generator: Automated Video Ad Workflow with Gemini 2.5 Pro

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Nutlope (Houssein Djirdeh) 使用 Llama 3.1 405B 处理真实任务执行

Nutlope (Houssein Djirdeh) · Llama 3.1 405B

A
厂商:Meta / Llama 模型:Llama 3.1 405B 来源平台:github 最后复核:2026-06-26T14:08:27Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Nutlope (Houssein Djirdeh) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T14:08:27Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LlamaCoder is an open-source clone of Claude Artifacts that generates small interactive web applications from a single natural-language prompt. Users describe an app idea and LlamaCoder produces a complete, runnable Nex…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully functional web app generator that produces working React/Next.js components from text prompts. The generated code is displayed and executed live in Sandpack sandboxes. The project was one of the most popular ope…

模型作用:Llama 3.1 405B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Llama 3.1 405B is the sole LLM powering the code generation pipeline. It receives user prompts describing desired mini-apps and generates complete, syntactically correct React/Next.js component code that runs in the bro…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Relies on Together AI's API for Llama 3.1 405B inference; API availability and pricing may change. Project is community-maintained, not an official Meta product.

原始记录:Nutlope LlamaCoder: Open-Source Claude Artifacts Clone Powered by Llama 3.1 405B

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SambaNova 使用 Llama 3.1 405B 处理知识检索和问答

SambaNova · Llama 3.1 405B

A
厂商:Meta / Llama 模型:Llama 3.1 405B 来源平台:web 最后复核:2026-06-26T13:08:46Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

SambaNova 公开的知识库与检索问答案例,来源为 web,复核于 2026-06-26T13:08:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SambaNova deployed Llama 3.1 405B on their Cloud platform optimized for function calling and agentic workflows, achieving up to 200 tokens/s. They demonstrated that while other providers' 405B instances average ~28 tok/…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production cloud service at cloud.sambanova.ai serving Llama 3.1 405B at 200 tokens/s with function calling capabilities. Demonstrated via AI Starter Kit with Streamlit UI and Jupyter notebook showing real-time multi-st…

模型作用:Llama 3.1 405B 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Llama 3.1 405B's function calling capabilities and instruction following precision make it the ideal model for agentic workflows where correct function inference and argument generation are critical. SambaNova's optimiz…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:SambaNova Cloud is a commercial service. Blog post dated September 2024. Pricing and availability may have changed.

原始记录:SambaNova Cloud Serves Llama 3.1 405B at 200 tok/s for Function Calling and Agentic Workflows

已有真实案例 知识库与检索问答webA 类可核验real_case auto_approved 进入模型卡精选

Nishit Verma (nishit-2611) 使用 Gemini 3.1 Pro Preview 处理知识检索和问答

Nishit Verma (nishit-2611) · Gemini 3.1 Pro Preview

A
厂商:Google / Gemini 模型:Gemini 3.1 Pro Preview 来源平台:GitHub 最后复核:2026-06-26T13:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Nishit Verma (nishit-2611) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T13:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Batch video captioning from Google Drive links using Gemini 3.1 Pro Preview. Users upload a CSV of Google Drive video URLs, configure system prompts, select model variant (gemini-3.1-pro-preview or customtools), set thi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Streamlit web application that: (1) reads CSV with Google Drive video links, (2) downloads shared videos, (3) uploads them to Gemini Files API, (4) generates detailed captions using gemini-3.1-pro-preview with configura…

模型作用:Gemini 3.1 Pro Preview 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3.1 Pro Preview is the core multimodal model for video understanding and caption generation. The pipeline specifically leverages Gemini 3.1 Pro's native video processing via the Files API, configurable thinking l…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Individual developer project. Requires Google Drive links shared with 'Anyone with the link can view' permissions. Deployable to Streamlit Community Cloud.

原始记录:Batch Video Captioning Pipeline for Google Drive Videos Using Gemini 3.1 Pro Preview

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

wanshuiyin (ARIS project) 使用 Gemini 3.1 Pro Preview 处理研究分析和报告生成

wanshuiyin (ARIS project) · Gemini 3.1 Pro Preview

A
厂商:Google / Gemini 模型:Gemini 3.1 Pro Preview 来源平台:GitHub 最后复核:2026-06-26T21:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

wanshuiyin (ARIS project) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ARIS (Auto-Research-In-Sleep) is an autonomous ML research system with 12.6k GitHub stars. Gemini 3.1 Pro (high reasoning) is configured as one of two primary models (alongside Claude Opus 4.6) for literature survey, cr…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production-grade autonomous research framework shipping 74+ user-facing skills and 49 helper resources. Gemini 3.1 Pro drives literature survey (/research-lit), idea discovery (/idea-discovery), and academic search (/…

模型作用:Gemini 3.1 Pro Preview 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3.1 Pro (high) is a primary reasoning model in the ARIS dual-model architecture. It handles large-context literature review, semantic search across academic databases, and participates in multi-model adversarial …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Large open-source project (12.6k stars). Gemini 3.1 Pro is one of multiple models used; not the sole model.

原始记录:ARIS: Autonomous ML Research Pipeline with Gemini 3.1 Pro as Core Model

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yeameen 使用 Gemini 3.1 Pro Preview 处理软件工程任务执行

yeameen · Gemini 3.1 Pro Preview

A
厂商:Google / Gemini 模型:Gemini 3.1 Pro Preview 来源平台:GitHub 最后复核:2026-06-26T21:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yeameen 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Claude Code plugin that runs seven AI reviewers in parallel on the same code diff, then synthesizes all findings. Gemini 3.1 Pro (via gemini CLI) is one of the seven reviewers — specifically tasked as a cross-mo…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Claude Code plugin (/plugin install review-council) that orchestrates 7 parallel reviewers: Codex CLI (GPT-5.5), Gemini CLI (Gemini 3.1 Pro), and 5 Claude specialist subagents. Produces synthesized reports with findin…

模型作用:Gemini 3.1 Pro Preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3.1 Pro provides independent cross-model perspective in the review council — the key insight is that different model families miss different things, so running Gemini alongside Claude and GPT catches blind spots.…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (16 stars). Gemini 3.1 Pro is one of three model families used in the council, not the sole reviewer.

原始记录:Multi-Agent Code Review Council with Gemini 3.1 Pro as Parallel Reviewer

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Amp (Sourcegraph / ampcode.com) 使用 Gemini 3 Pro Preview (high) 处理软件工程任务执行

Amp (Sourcegraph / ampcode.com) · Gemini 3 Pro Preview (high)

A
厂商:Google / Gemini 模型:Gemini 3 Pro Preview (high) 来源平台:web 最后复核:2026-06-26T21:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Amp (Sourcegraph / ampcode.com) 公开的代码代理与软件工程案例,来源为 web,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Amp switched its primary smart-agent-mode model from Claude to Gemini 3 Pro, using it as the default for all code-editing, refactoring, and agentic coding tasks in their open-source AI coding assistant.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Gemini 3 Pro became the main model powering Amp's smart agent mode. The team reported 'ecstatic messages' in their Slack, with users calling it 'crazy, this is incredible' and noting its superior balance of intelligence…

模型作用:Gemini 3 Pro Preview (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3 Pro provided the best balance of intelligence, speed, and tool-use capability among all models evaluated (including GPT-5, Gemini 2.5, Grok). It replaced Claude as the default, marking the first model change si…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Source is official product blog from ampcode.com. Model name is 'Gemini 3 Pro' (marketing name); the API variant gemini-3-pro-preview-high corresponds to the high thinking-budget tier.

原始记录:Amp adopts Gemini 3 Pro as default smart agent model, replacing Claude

已有真实案例 代码代理与软件工程webA 类可核验real_case auto_approved 进入模型卡精选

Simon Willison 使用 Gemini 3 Pro Preview (high) 处理代码审查和测试生成

Simon Willison · Gemini 3 Pro Preview (high)

A
厂商:Google / Gemini 模型:Gemini 3 Pro Preview (high) 来源平台:web 最后复核:2026-06-26T21:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Simon Willison 公开的代码审查与测试案例,来源为 web,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Simon Willison used Gemini 3 Pro's multimodal audio capabilities via the LLM CLI tool to transcribe a 3-hour 33-minute Half Moon Bay City Council meeting audio file (74MB m4a), requesting a Markdown transcript with spea…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Gemini 3 Pro successfully transcribed the full 3h33m meeting into structured Markdown with speaker identification and timestamps. Willison also ran a 'pelican benchmark' testing Gemini 3 Pro's vision capabilities across…

模型作用:Gemini 3 Pro Preview (high) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3 Pro's native multimodal audio processing handled a 74MB audio file end-to-end, producing a structured transcript in a single API call — demonstrating production-quality long-form audio transcription capability …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Source is Simon Willison's well-known technical blog. The exact API model used was 'gemini-3-pro-preview' (confirmed in the blog post command). The '(high)' tier corresponds to the high thinking budget setting.

原始记录:Simon Willison transcribes 3.5-hour city council meeting with Gemini 3 Pro audio

已有真实案例 代码审查与测试webA 类可核验real_case auto_approved 进入模型卡精选

Joel Zhang (blog.jcz.dev / Gemini P… 使用 Gemini 3 Pro Preview (high) 处理可玩交互原型构建

Joel Zhang (blog.jcz.dev / Gemini Plays Pokemon project) · Gemini 3 Pro Preview (high)

A
厂商:Google / Gemini 模型:Gemini 3 Pro Preview (high) 来源平台:github 最后复核:2026-06-26T21:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Joel Zhang (blog.jcz.dev / Gemini Plays Pokemon project) 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Joel Zhang ran Gemini 3 Pro inside the Gemini Plays Pokemon harness — an autonomous agent framework with tools for mental mapping, inventory management, and game interaction — to play Pokemon Crystal from start to finis…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Gemini 3 Pro became the Johto Champion without losing a single battle, using approximately half as many turns as Gemini 2.5 Pro. The author described it as 'a different species of agent' that could ground its knowledge …

模型作用:Gemini 3 Pro Preview (high) 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3 Pro demonstrated superior visual grounding and strategic reasoning by using ~50% fewer turns than Gemini 2.5 Pro, successfully interpreting game screen pixels to make informed decisions, avoid soft-locks, and c…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Source is a detailed personal blog post with direct API usage evidence. The harness (Gemini Plays Pokemon) is open-source. Model used was the Gemini 3 Pro preview variant.

原始记录:Gemini 3 Pro completes Pokemon Crystal without losing a single battle

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Gordon Pimblott (bitwrangler.uk) 使用 Gemini 3 Pro Preview (high) 处理软件工程任务执行

Gordon Pimblott (bitwrangler.uk) · Gemini 3 Pro Preview (high)

A
厂商:Google / Gemini 模型:Gemini 3 Pro Preview (high) 来源平台:web 最后复核:2026-06-26T21:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Gordon Pimblott (bitwrangler.uk) 公开的代码代理与软件工程案例,来源为 web,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Gordon Pimblott dusted off a shelved C++ ZX Spectrum emulator project (abandoned 2 years prior due to the grind of implementing hundreds of Z80 opcodes) and used Gemini 3 Pro via Google Antigravity IDE to complete the i…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The emulator was booting the BASIC ROM in a single evening, whereas the original project had been abandoned after weeks of manual work. Gemini 3 Pro handled boilerplate opcode implementations and logic code effectively,…

模型作用:Gemini 3 Pro Preview (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3 Pro acted as a 'powerful accelerator' for boilerplate and logic implementation, completing in one evening what would have taken weeks of manual Z80 opcode coding. Its agentic workflow in Antigravity (creating .…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Source is a personal blog post with detailed account. The developer used Gemini 3 Pro through Google Antigravity IDE (not direct API), but explicitly named the model. The post includes honest assessment of both successe…

原始记录:Developer finishes C++ ZX Spectrum emulator in one evening using Gemini 3 Pro

已有真实案例 代码代理与软件工程webA 类可核验real_case auto_approved 进入模型卡精选

letmutex (GitHub) 使用 Gemini 3 Pro Preview (high) 处理软件工程任务执行

letmutex (GitHub) · Gemini 3 Pro Preview (high)

A
厂商:Google / Gemini 模型:Gemini 3 Pro Preview (high) 来源平台:github 最后复核:2026-06-26T21:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

letmutex (GitHub) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T21:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A developer used Gemini 3 Pro to code a Python implementation of the Line Girl visual effect (a variant of the Joy Division album cover effect), creating an animated visualization from image data.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Gemini 3 Pro generated a working Python implementation of the Line Girl effect as a complete, runnable project with video output. The developer's README explicitly states 'Coded by Gemini 3 Pro and tweaked by me,' with …

模型作用:Gemini 3 Pro Preview (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3 Pro generated the entire initial codebase for the visual effect — including the mathematical transformations, rendering pipeline, and animation logic — which the developer then tweaked for final polish. The pro…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Source is a GitHub repo with explicit attribution in README. The HN discussion (item 46288176) confirms the project was generated by Gemini 3 Pro. Model name is 'Gemini 3 Pro' (marketing name, maps to gemini-3-pro-previ…

原始记录:Developer codes Line Girl visual effect in Python using Gemini 3 Pro

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Omni AI (getomni-ai) 使用 Claude 3 Haiku 处理知识检索和问答

Omni AI (getomni-ai) · Claude 3 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3 Haiku 来源平台:github 最后复核:2026-06-26T13:17:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Omni AI (getomni-ai) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T13:17:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Zerox is a document OCR library (12,243 GitHub stars) that converts PDFs, DOCX files, and images into Markdown by passing pages as images to vision models. It explicitly lists Claude 3 Haiku (2024.03, 2024.10) as a supp…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Structured Markdown output from PDF/image documents; used for downstream LLM ingestion, data extraction pipelines, and document search systems

模型作用:Claude 3 Haiku 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3 Haiku provides fast, cost-effective vision-based OCR as one of Zerox's Bedrock-supported models. At $0.25/$1.25 per million tokens (input/output), it enables processing large document volumes at a fraction of t…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Zerox defaults to OpenAI GPT-4o for OCR; Claude 3 Haiku is one of several supported providers. Users must explicitly configure the Bedrock/Haiku provider. The product supports Anthropic direct API and AWS Bedrock, but t…

原始记录:Zerox OCR: Vision-based Document Extraction Using Claude 3 Haiku via AWS Bedrock

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Lemony AI (lemony-ai) 使用 Claude 3 Haiku 处理智能体流程编排

Lemony AI (lemony-ai) · Claude 3 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3 Haiku 来源平台:github 最后复核:2026-06-26T13:17:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Lemony AI (lemony-ai) 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T13:17:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:CascadeFlow is an agent runtime intelligence layer (2,740 GitHub stars) that dynamically selects the optimal LLM for each query or tool call through speculative execution. It uses Claude 3 Haiku as the initial speculati…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Per-query optimal model routing decisions; 40-85% cost reduction while retaining 96% quality of flagship models; sub-5ms routing overhead in agent execution loops

模型作用:Claude 3 Haiku 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3 Haiku serves as the primary speculative execution model, handling the majority of queries at $0.15-0.30/1M tokens before potential escalation to more expensive models. This makes it the cost-efficiency backbone…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:CascadeFlow is a framework/tool that uses Haiku as part of its multi-model cascading strategy, not a standalone Haiku product deployment. The actual production impact depends on individual users' configurations and task…

原始记录:CascadeFlow: Claude 3 Haiku as Speculative Execution Model in Agent Cascading Runtime

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

vxcontrol 使用 Claude 3 Haiku 处理代码审查和测试生成

vxcontrol · Claude 3 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3 Haiku 来源平台:github 最后复核:2026-06-26T13:17:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

vxcontrol 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T13:17:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:PentAGI is a fully autonomous penetration testing platform (17,973 GitHub stars) that uses AI agents to perform complex security assessment tasks. It explicitly lists claude-3-haiku-20240307 as a supported LLM provider …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Comprehensive vulnerability reports with exploitation guides; automated penetration test execution across network reconnaissance, web application testing, and privilege escalation phases

模型作用:Claude 3 Haiku 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3 Haiku provides fast, cost-effective AI reasoning for real-time security analysis, automated task planning, and intelligent tool selection during penetration testing. Its speed (21K tokens/second) enables rapid …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:PentAGI supports 10+ LLM providers; Claude 3 Haiku is one option among many. Users can configure any supported provider. The pentagi-3-haiku-20240307 model is noted as being retired April 19, 2026, with migration to cla…

原始记录:PentAGI: Claude 3 Haiku for Automated Penetration Testing AI Agent Reasoning

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Tamer Comcuoglu 使用 GPT-4o (May) 处理研究分析和报告生成

Tamer Comcuoglu · GPT-4o (May)

A
厂商:OpenAI 模型:GPT-4o (May) 来源平台:personal_blog 最后复核:2026-06-26T21:50:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Tamer Comcuoglu 公开的研究与报告生成案例,来源为 博客记录,复核于 2026-06-26T21:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Tamer Comcuoglu used GPT-4o with LangChain's structured output (JSON mode) to classify and extract structured fields from over 10,000 'Ask HN: Who is Hiring' comments spanning May 2022 to June 2024. Each comment was par…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Successfully classified 10,891 job posting comments into structured records covering ~$54 in GPT-4o API costs. Produced actionable labor market insights: remote work trends, visa sponsorship rates, experience level dist…

模型作用:GPT-4o (May) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o was the core engine for structured data extraction, using JSON mode to output typed fields from unstructured HTML comment text. Its instruction-following capability enabled reliable field extraction including boo…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Blog post and HN discussion (465 points) are publicly accessible. Author is an individual developer, not a company. The GPT-4o model used is explicitly named. API costs ($54) are documented.

原始记录:HN Job Market Insights: Structured Analysis of 10,000+ Comments Using GPT-4o

已有真实案例 研究与报告生成博客记录A 类可核验real_case auto_approved 进入模型卡精选

Yigit Konur 使用 GPT-4o (May) 处理文档理解和结构化处理

Yigit Konur · GPT-4o (May)

A
厂商:OpenAI 模型:GPT-4o (May) 来源平台:github 最后复核:2026-06-26T21:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Yigit Konur 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T21:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Yigit Konur built api-llm-ocr, an open-source FastAPI service that uses GPT-4o's vision API to convert PDF documents into clean, structured markdown. The tool handles complex documents with tables, mixed layouts, header…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production-ready tool that processes PDFs at ~1,500 tokens per page using GPT-4o vision. Cost: ~$15 per 1,000 pages with GPT-4o, ~$8 with GPT-4o mini, ~$4 with batch API. Supports parallel processing (50 pages in seco…

模型作用:GPT-4o (May) 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o's multimodal vision capability is the core differentiator — it reads and understands document context, not just character shapes like traditional OCR. The model correctly identifies table structures, headers, and…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source GitHub repo (AGPL v3 license). GPT-4o explicitly listed in the cost comparison table. Demo video available in the repo. Author is an independent developer. Repo includes working code with configuration for O…

原始记录:api-llm-ocr: Production PDF-to-Markdown Conversion Using GPT-4o Vision

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Matthieu Le Cauchois 使用 GPT-4o (May) 处理多模态内容处理

Matthieu Le Cauchois · GPT-4o (May)

A
厂商:OpenAI 模型:GPT-4o (May) 来源平台:personal_blog 最后复核:2026-06-26T21:50:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Matthieu Le Cauchois 公开的多模态生成与理解案例,来源为 博客记录,复核于 2026-06-26T21:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Matthieu Le Cauchois built Shoggoth Mini, a soft tentacle robot that uses GPT-4o's real-time API for voice-driven interaction and high-level behavioral control. GPT-4o continuously listens to speech through the audio st…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A functional physical robot with expressive behaviors driven by GPT-4o's real-time audio API. The system has two control layers: low-level (hardcoded primitives + RL policies) and high-level (GPT-4o decision-making). GP…

模型作用:GPT-4o (May) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o's real-time API is the high-level brain of the system — it continuously processes the audio stream for speech recognition and makes contextual decisions about robot behavior. The model's ability to understand nat…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Blog post includes detailed technical documentation and video demonstrations of the robot in action. Author is a robotics engineer. The project is described as a hobby/research project, not a commercial product. GPT-4o …

原始记录:Shoggoth Mini: Expressive Soft Tentacle Robot with GPT-4o Real-Time Voice Control

已有真实案例 多模态生成与理解博客记录A 类可核验real_case auto_approved 进入模型卡精选

Eduardo Blancas 使用 GPT-4o (May) 处理医疗和生命科学分析

Eduardo Blancas · GPT-4o (May)

A
厂商:OpenAI 模型:GPT-4o (May) 来源平台:personal_blog 最后复核:2026-06-26T21:50:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Eduardo Blancas 公开的医疗与生命科学案例,来源为 博客记录,复核于 2026-06-26T21:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕医疗和生命科学分析的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Eduardo Blancas built an AI-assisted web scraper that uses GPT-4o's structured output feature to extract structured data from HTML tables and complex web pages. The system can parse complex table layouts (merged rows, n…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Successfully extracted structured data from complex HTML tables including: 10-day weather forecasts with nested day/night rows, Wikipedia tables with merged cells (Human Development Index), and hidden HTML elements that…

模型作用:GPT-4o (May) 在该案例中承担医疗和生命科学分析相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o's structured output feature (JSON mode) is the core of the system — it parses raw HTML into typed, schema-validated data structures. The model's vision-like understanding of HTML structure enables it to handle co…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Blog post includes demo links and source code references. Author is a data scientist (blancas.io). The project demonstrates both the power and cost limitations of GPT-4o for web scraping (377pts HN discussion). Blog pos…

原始记录:AI-Assisted Web Scraper: Structured Data Extraction and XPath Generation with GPT-4o

已有真实案例 医疗与生命科学博客记录A 类可核验real_case auto_approved 进入模型卡精选

Olup (product engineering team) 使用 GPT-4o (May) 处理多模态内容处理

Olup (product engineering team) · GPT-4o (May)

A
厂商:OpenAI 模型:GPT-4o (May) 来源平台:personal_blog 最后复核:2026-06-26T21:50:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Olup (product engineering team) 公开的多模态生成与理解案例,来源为 博客记录,复核于 2026-06-26T21:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Olup (a product engineer) used GPT-4o for production image detection to distinguish between 350 highly similar car illustrations in a museum setting. Users photograph an illustration on a wall, and the system must ident…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A deployed production system that accurately identifies which of 350 similar car illustrations a user has photographed. The pipeline combines a fast embedding-based KNN filter with GPT-4o for final disambiguation, balan…

模型作用:GPT-4o (May) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o serves as the final disambiguation step in a multi-stage vision pipeline. After KNN filtering narrows candidates from 350 to a handful, GPT-4o's vision capability performs fine-grained visual comparison to identi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Blog post is on a Pages.dev hosted site (may be less permanent). HN discussion (222 points) confirms it is a real production deployment. The team is described as a 'crafty product engineer team' — likely a small agency …

原始记录:Production Image Detection for 350 Similar Museum Illustrations Using GPT-4o

已有真实案例 多模态生成与理解博客记录A 类可核验real_case auto_approved 进入模型卡精选

Shubhamsaboo 使用 Gemini 3.1 Pro Preview 处理真实任务执行

Shubhamsaboo · Gemini 3.1 Pro Preview

A
厂商:Google / Gemini 模型:Gemini 3.1 Pro Preview 来源平台:GitHub 最后复核:2026-06-26T13:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Shubhamsaboo 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T13:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A self-contained creative coding playground where users describe a visual concept in natural language and Gemini 3.1 Pro generates website-ready, animated canvas/SVG code in real-time at 60fps. Supports iterative refine…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Single-HTML-file creative coding tool (no dependencies, no build step) with streaming SSE code generation, model selector defaulting to gemini-3.1-pro-preview, preset prompt chips (particle galaxies, Matrix rain, fracta…

模型作用:Gemini 3.1 Pro Preview 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3.1 Pro Preview is the default and recommended model ('deep reasoning, intricate simulations'). The entire tool architecture — streaming code injection, 60fps animation loop, error-recovery prompt chain — is desi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Individual developer open-source project (MIT license), not a large-scale production deployment. 71 GitHub stars as of collection date.

原始记录:Gemini Artboard — Real-Time Creative Coding Playground Powered by Gemini 3.1 Pro

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Princeton NLP / SWE-agent (Karol Ra… 使用 Claude 3.7 Sonnet 处理软件工程任务执行

Princeton NLP / SWE-agent (Karol Rajewski et al.) · Claude 3.7 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 3.7 Sonnet 来源平台:github 最后复核:2026-06-26T13:55:38Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Princeton NLP / SWE-agent (Karol Rajewski et al.) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:55:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The SWE-agent team at Princeton NLP used Claude 3.7 Sonnet as the backbone LLM for their open-source autonomous software engineering agent. SWE-agent 1.0 + Claude 3.7 achieved state-of-the-art results on both SWE-bench …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:SWE-agent 1.0 paired with Claude 3.7 Sonnet achieved the highest resolve rates on SWE-bench Verified and Full leaderboards at the time of release, surpassing all prior agent systems. The source code explicitly handles C…

模型作用:Claude 3.7 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.7 Sonnet served as the reasoning backbone for SWE-agent's autonomous software engineering pipeline. Its code understanding, multi-step reasoning, and tool-use capabilities enabled the agent to autonomously navi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:SWE-bench results are benchmark evaluations, but SWE-agent is a real open-source tool used for autonomous software engineering on real codebases. The model's contribution is demonstrated through a concrete, reproducible…

原始记录:SWE-agent 1.0 achieves state-of-the-art on SWE-bench using Claude 3.7 Sonnet as backbone model

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Aider-AI 使用 o1 处理软件工程任务执行

Aider-AI · o1

A
厂商:OpenAI 模型:o1 来源平台:github 最后复核:2026-06-27T07:08:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Aider-AI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:08:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Aider is an open-source command-line pair-programming tool that maps a codebase, edits files, supports many programming languages, and commits AI-generated changes with git integration.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository and documentation describe repository-aware code editing, codebase maps, multi-language support, and automatic git commits; the README lists OpenAI o1 as a recommended supported model.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI o1 supplies the reasoning/coding model option for complex code changes and repository-level programming assistance inside the Aider workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is product documentation for model support in a public coding artifact, not a single named downstream customer deployment. It is included because the organization, task, artifact, and o1 binding are explicit.

原始记录:Aider supports OpenAI o1 for repository-aware pair programming

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

NVIDIA 使用 Llama 3.1 405B 处理文档理解和结构化处理

NVIDIA · Llama 3.1 405B

A
厂商:Meta / Llama 模型:Llama 3.1 405B 来源平台:github 最后复核:2026-06-26T14:08:27Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

NVIDIA 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T14:08:27Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:NVIDIA published an official AI Blueprint that transforms PDF documents into engaging audio podcasts. The system uses a 3-LLM ensemble consisting of Llama 3.1-8B, Llama 3.1-70B, and Llama 3.1-405B NIMs to handle differe…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production-ready pipeline that ingests PDF documents and outputs multi-speaker podcast audio files. The ensemble approach uses the 405B model for high-quality script generation and complex content understanding, while…

模型作用:Llama 3.1 405B 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Llama 3.1 405B is the flagship model in the 3-model ensemble, responsible for the most cognitively demanding tasks: deep content comprehension of PDF documents, generation of engaging conversational scripts, and maintai…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Requires NVIDIA NIM infrastructure and GPU access. Blueprint is an NVIDIA reference architecture, not a standalone product. Llama 3.1 405B NIM may require specific hardware (H100/A100).

原始记录:NVIDIA AI Blueprint: PDF-to-Podcast with Llama 3.1 405B Ensemble

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Raindrop.ai (Ben Hylak, co-founder … 使用 o3-pro 处理真实任务执行

Raindrop.ai (Ben Hylak, co-founder & Alexis Gauba, co-founder) · o3-pro

A
厂商:OpenAI 模型:o3-pro 来源平台:latent.space 最后复核:2026-06-26T22:30:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Raindrop.ai (Ben Hylak, co-founder & Alexis Gauba, co-founder) 公开的真实任务执行案例,来源为 latent.space,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Raindrop.ai 联合创始人 Ben Hylak 和 Alexis Gauba 将公司历史上的所有规划会议记录、目标文档和语音备忘录汇总后,交给 o3-pro 进行深度分析,要求生成一份包含目标指标、时间表、优先级排序和明确裁减建议的完整战略规划。

公开产物:o3-pro 生成了极其具体且可执行的战略计划,包含精确的目标指标、时间线、优先级排序和明确的取舍建议。Ben Hylak 在文中表示:'The plan o3 Pro gave us was specific and rooted enough that it actually changed how we are thinking about our future.'(o3 Pro 给出的计划足够具体和扎实,真正改变了我们对未来的思考方式。)

模型作用:o3-pro 在此任务中的核心贡献是其超长上下文处理能力和深度推理能力。它能够消化大量异构历史数据(会议记录、目标、语音备忘录),并生成结构化的战略分析,这是 o3 基础模型无法达到的深度。文章指出 o3-pro 'smarter, much smarter'(聪明得多),但需要大量上下文输入才能发挥优势。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该案例来自 Latent Space 博客的一篇早期评测文章,作者 Ben Hylak 是 Raindrop.ai 联合创始人,属于第一方叙述。Raindrop.ai 是一家真实的 AI Agent 监控公司(https://www.raindrop.ai),有 Fortune 100 客户和 857 GitHub stars。案例描述了具体的使用场景和可验证的结果,但未提供 o3-pro 生成的具体规划文档作为独立 artifact。

原始记录:Raindrop.ai 使用 o3-pro 进行公司战略规划

已有真实案例 真实任务执行latent.spaceA 类可核验real_case auto_approved 进入模型卡精选

Roboflow (authored by James Gallagh… 使用 o3-pro 处理代码审查和测试生成

Roboflow (authored by James Gallagher) · o3-pro

A
厂商:OpenAI 模型:o3-pro 来源平台:blog 最后复核:2026-06-26T14:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Roboflow (authored by James Gallagher) 公开的代码审查与测试案例,来源为 博客记录,复核于 2026-06-26T14:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Roboflow, a computer vision platform company, evaluated OpenAI o3-pro across multiple vision use cases including industrial defect detection, object counting in complex scenes, and Visual Question Answering (VQA) to ass…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Roboflow published detailed evaluation results showing o3-pro's performance on defect detection, object counting, and VQA tasks, with comparisons to other frontier models. The article demonstrates o3-pro's strengths and…

模型作用:o3-pro 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:o3-pro provided multimodal reasoning capability that enabled complex visual analysis tasks including defect detection, counting, and question answering about images. The evaluation showed it as a competitive option for …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The article is a review/evaluation by Roboflow rather than a deployment story. Roboflow is a real company with a published product. Article last modified 2026-06-23, indicating ongoing updates. Content behind Cloudflare…

原始记录:Roboflow evaluates o3-pro for industrial computer vision tasks

已有真实案例 代码审查与测试博客记录A 类可核验real_case auto_approved 进入模型卡精选

Simon Willison (developer, Datasett… 使用 o3-pro 处理真实任务执行

Simon Willison (developer, Datasette creator) · o3-pro

A
厂商:OpenAI 模型:o3-pro 来源平台:blog 最后复核:2026-06-26T14:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Simon Willison (developer, Datasette creator) 公开的真实任务执行案例,来源为 博客记录,复核于 2026-06-26T14:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Simon Willison tested o3-pro on launch day by using his llm-openai-plugin to generate an SVG of a pelican riding a bicycle via the command line. He also noted that o3-pro works best when combined with tools and tracked …

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:o3-pro successfully generated the SVG of a pelican riding a bicycle, though it took 124 seconds. Willison noted o3-pro is priced at $20/M input tokens and $80/M output tokens (10x o3 after its 80% price drop). He highli…

模型作用:o3-pro 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o3-pro demonstrated its reasoning capability for code generation tasks, producing functional SVG output. The latency (124 seconds for a single SVG) highlighted the compute-intensive nature of the model. Willison's integ…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a developer testing/demo scenario rather than a production deployment. However, Simon Willison is a highly credible, named developer with a public track record. The blog post is freely accessible and documents s…

原始记录:Simon Willison uses o3-pro via llm CLI for SVG generation

已有真实案例 真实任务执行博客记录A 类可核验real_case auto_approved 进入模型卡精选

mikedemarais 使用 Grok 4 处理软件工程任务执行

mikedemarais · Grok 4

A
厂商:xAI / Grok 模型:Grok 4 来源平台:github 最后复核:2026-06-26T13:45:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mikedemarais 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:45:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Claude Code Skill that routes X/Twitter queries to xAI Grok 4 via OpenRouter with Live Search enabled, scoped to the X source. Supports handle filtering, date range queries, engagement filters (favorites/views), and c…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:JSON responses with summary of X/Twitter results, tweet URL citations, and usage stats. The skill provides real-time social media intelligence by leveraging Grok 4's Live Search capability.

模型作用:Grok 4 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 4 is the sole inference model used for all query processing and summarization. The default model is explicitly set to 'x-ai/grok-4' in the source code. Grok 4's Live Search feature is the core capability that enabl…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small repo (18 stars). Model accessed via OpenRouter, not directly via xAI API. Grok 4 model ID confirmed as default in source: const model = process.env.GROK_MODEL ?? 'x-ai/grok-4'.

原始记录:Claude Code Skill for Real-Time X/Twitter Search via Grok 4

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MORNLONG 使用 Grok 4 处理研究分析和报告生成

MORNLONG · Grok 4

A
厂商:xAI / Grok 模型:Grok 4 来源平台:github 最后复核:2026-06-26T13:45:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MORNLONG 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T13:45:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:An AI-powered cryptocurrency automated trading bot integrating xAI Grok-4 with OKX exchange API. Provides 24/7 market analysis using multi-timeframe technical indicators (50+ indicators), volume profile analysis, and st…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Structured trading decisions including buy/sell signals, position management, risk management (stop-profit/stop-loss), and automated execution via OKX API. Supports configurable execution intervals and real-time market …

模型作用:Grok 4 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Grok-4 is the core AI model for all market analysis and trading decision generation. It processes multi-timeframe technical indicators and generates structured output for reliable automated trading decisions. The README…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small repo (6 stars). Trading involves financial risk - bot explicitly warns about potential total fund loss. Docker support available. Model binding confirmed in README: 'xAI Grok-4 Integration'.

原始记录:AI-Powered Cryptocurrency Trading Bot with Grok-4

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DevDizzle 使用 Grok 4 处理文档理解和结构化处理

DevDizzle · Grok 4

A
厂商:xAI / Grok 模型:Grok 4 来源平台:github 最后复核:2026-06-26T13:45:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DevDizzle 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T13:45:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A local multi-agent invoice processing system built with LangGraph. The pipeline: Ingestion (OCR/LLM extraction) → Validation (line items checked against local SQLite inventory) → Approval (VP Agent scrutinizes high-val…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Structured invoice processing pipeline output: extracted line items, validation results against inventory, VP approval/rejection decisions with reasoning, and payment recommendations. Test scenarios include standard app…

模型作用:Grok 4 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Grok 4.1 Fast Reasoning (a grok-4 family variant) powers the intelligent data extraction from invoices (OCR/LLM) and the reflective VP approval agent that scrutinizes high-value invoices. Model binding confirmed in READ…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small repo (2 stars). Uses Grok 4.1 Fast Reasoning variant, not the base grok-4. Deployed on Google Cloud Run (profitscout-lx6bb project). XAI_API_KEY required.

原始记录:Autonomous Multi-Agent Invoice Processing System

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LMSYS / SGLang Project 使用 MiMo-V2-Flash (Feb 2026) 处理多模态内容处理

LMSYS / SGLang Project · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T14:10:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LMSYS / SGLang Project 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SGLang team (contributor acelyc111) integrated MiMo-V2-Flash as a first-class supported model in the SGLang LLM serving framework, implementing optimized Sliding Window Attention (SWA) execution, multi-layer MTP (Multi-…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Merged PR #15207 with full day-0 support. Achieved 150 TPS per request decoding throughput even under 64K input tokens with batch size 16 per DP rank. Published detailed benchmarking results for prefill and decode phase…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash's hybrid SWA+GA attention architecture and 3-layer MTP design enabled SGLang to demonstrate that inference-centric model design can achieve balanced throughput and latency. The model's architecture was spe…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:MiMo-V2-Flash was retired from Xiaomi MiMo Open Platform on 2026-06-18; model may no longer be available via official API. Historical use case.

原始记录:SGLang Day-0 Support for MiMo-V2-Flash with Optimized Inference

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

aaronjmars (MiroShark project) 使用 MiMo-V2-Flash (Feb 2026) 处理文档理解和结构化处理

aaronjmars (MiroShark project) · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T14:10:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

aaronjmars (MiroShark project) 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MiroShark (1,337 stars), a Universal Swarm Intelligence Engine for financial forecasting and future prediction, used xiaomi/mimo-v2-flash as its default LLM model via OpenRouter for persona generation, simulation config…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production deployment as default model across multiple simulation slots (Default + Wonderwall). Each simulation run consumed 7M+ tokens across 850+ agent-action calls. The project documented MiMo-V2-Flash pricing at $0.…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash's combination of strong reasoning capability and competitive pricing ($0.10/M input) made it viable as the default model for high-volume swarm intelligence simulations. The model's persona generation quali…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:MiMo-V2-Flash was delisted from OpenRouter on 2026-06-22; project switched to MiMo-V2.5. Historical use case.

原始记录:MiroShark: MiMo-V2-Flash as Default Model for Swarm Intelligence Simulations

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Alibaba Group / OpenCodeReview 使用 MiMo-V2-Flash (Feb 2026) 处理代码审查和测试生成

Alibaba Group / OpenCodeReview · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T14:10:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Alibaba Group / OpenCodeReview 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Alibaba's OpenCodeReview (9,380 stars), an AI-powered code review CLI tool battle-tested at Alibaba's scale serving tens of thousands of developers, included MiMo-V2-Flash as a supported model in its built-in MiMo provi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:MiMo-V2-Flash was listed as a supported model in OpenCodeReview's built-in provider registry alongside other MiMo models (mimo-v2-pro, mimo-v2-omni). The model was removed in PR #192 (merged 2026-06-23) after Xiaomi ann…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash served as one of the recommended models for AI-powered code review in Alibaba's OpenCodeReview tool. Its reasoning and coding capabilities made it suitable for the tool's hybrid architecture combining dete…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:MiMo-V2-Flash deprecated from Xiaomi platform on 2026-06-30; removed from OpenCodeReview provider registry. Historical use case.

原始记录:Alibaba OpenCodeReview: MiMo-V2-Flash as Supported LLM for AI Code Review

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

theopenco (LLMGateway project) 使用 MiMo-V2-Flash (Feb 2026) 处理真实任务执行

theopenco (LLMGateway project) · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T14:10:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

theopenco (LLMGateway project) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LLMGateway (1,338 stars), a unified API interface for routing, managing, and analyzing LLM requests across multiple providers, had MiMo-V2-Flash configured as an active model route through the Xiaomi MiMo Open Platform …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:MiMo-V2-Flash was an active routed model in LLMGateway's provider registry. When Xiaomi announced the model would be retired on 2026-06-18 (with silent auto-forwarding to MiMo-V2.5), LLMGateway proactively deactivated t…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash was integrated as a production-ready model option in LLMGateway's multi-provider routing system, demonstrating its compatibility with standard OpenAI-compatible API interfaces and its viability as a routin…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:MiMo-V2-Flash retired from Xiaomi MiMo Open Platform on 2026-06-18; model route deactivated. Historical use case.

原始记录:LLMGateway: MiMo-V2-Flash as Routed Model in Multi-Provider LLM Gateway

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Factory-AI 使用 MiMo-V2-Pro 处理软件工程任务执行

Factory-AI · MiMo-V2-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Pro 来源平台:github 最后复核:2026-06-26T13:59:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Factory-AI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Factory-AI integrated MiMo-V2-Pro as a custom provider in their AI coding agent to autonomously generate, review, and refactor code. The agent uses MiMo-V2-Pro's reasoning/thinking capabilities for complex software engi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Factory AI coding agent running on MiMo-V2-Pro as its backend model, generating code with reasoning traces. The team filed a bug (issue #885) about thinking content being exposed in output, confirming active production …

模型作用:MiMo-V2-Pro 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Pro provides the core reasoning engine for the Factory coding agent, enabling chain-of-thought code generation and complex multi-file refactoring tasks. The model's 1M context window supports large codebase unde…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Bug report indicates thinking content leaks into output; issue filed June 2026 confirms active usage.

原始记录:Factory AI Coding Agent uses MiMo-V2-Pro for autonomous code generation

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Xiaomi MiMo Team 使用 MiMo-V2-Pro 处理多模态内容处理

Xiaomi MiMo Team · MiMo-V2-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Pro 来源平台:github 最后复核:2026-06-26T13:59:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Xiaomi MiMo Team 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T13:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Xiaomi built MiMo-Code as a terminal-native AI coding assistant that reads/writes code, executes commands, manages Git, and maintains persistent memory across sessions. MiMo-V2-Pro is the default reasoning model behind …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:MiMo-Code product with 605+ GitHub issues from active users, supporting multiple agents (build/plan/compose), persistent memory via SQLite FTS5, task tracking, subagent system, voice input via MiMo ASR, and Max Mode (pa…

模型作用:MiMo-V2-Pro 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Pro serves as the core intelligence engine for MiMo-Code's build agent, powering code generation, file editing, command execution, and long-horizon reasoning. The model's 1M context window enables deep project u…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:605+ issues on GitHub indicate broad adoption. Many issues report loop-thinking and token consumption concerns with mimo-v2-pro.

原始记录:MiMo-Code: Xiaomi's terminal-native AI coding assistant powered by MiMo-V2-Pro

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

alrobles 使用 MiMo-V2-Pro 处理医疗和生命科学分析

alrobles · MiMo-V2-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Pro 来源平台:github 最后复核:2026-06-26T13:59:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

alrobles 公开的医疗与生命科学案例,来源为 公开代码库,复核于 2026-06-26T13:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕医疗和生命科学分析的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:alrobles built ecoSeek Kids as a kid-friendly (ages 6-15) scientific chatbot focused on ecology, animals, plants, earth science, space, and human body topics. The backend routes all chat through a Hermes Gateway which c…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Live production chatbot at kids.ecoseek.org served via Cloudflare Tunnel. Features an animated mascot (Emily Astronauta), topic-specific responses for 6 science domains, and GBIF biodiversity data enrichment. Python ser…

模型作用:MiMo-V2-Pro 在该案例中承担医疗和生命科学分析相关的生成、分析、编排或实现角色。 原始资料写作:MiMo V2 Pro (V2.5 Pro variant) provides the reasoning backbone for science Q&A, enabling the chatbot to generate age-appropriate, factually grounded ecology and nature education content for children. Temperature set to …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Uses MiMo V2.5 Pro (closely related variant in same V2 family). GitHub issue #9 reports API quota exhaustion (429), confirming active production usage.

原始记录:ecoSeek Kids: Kid-friendly ecology study chatbot powered by MiMo V2 Pro

已有真实案例 医疗与生命科学公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Miku-cy 使用 MiMo-V2-Pro 处理软件工程任务执行

Miku-cy · MiMo-V2-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Pro 来源平台:github 最后复核:2026-06-26T13:59:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Miku-cy 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T13:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Miku-cy built a compatibility fix tool (API proxy + source patch + config fix) to solve MiMo API's 400 error when assistant messages contain tool_calls without reasoning_content field. The tool enables MiMo-V2-Pro to wo…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:MIT-licensed v1.1.0 tool with API proxy (auto-fills reasoning_content), OpenClaw source patch, and config fixes for 8 OpenAI-protocol and 7 Anthropic-protocol tools. Includes reasoning_content caching for multi-turn con…

模型作用:MiMo-V2-Pro 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Pro's reasoning/thinking capability is the core value proposition driving this compatibility tool. The 1M context and strong coding performance make it worth the effort to fix API compatibility across 15+ coding…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Tool was created specifically to fix MiMo API limitations. The need for this tool indicates real-world friction but also strong user demand for MiMo-V2-Pro.

原始记录:MiMo-API-400-Fix: Community compatibility tool enabling MiMo-V2-Pro use with Cursor, Claude Code, and 15+ tools

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CherryHQ (CherryHQ/cherry-studio) 使用 Seed-2.1-Pro-Preview 处理多模态内容处理

CherryHQ (CherryHQ/cherry-studio) · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:github 最后复核:2026-06-26T16:37:40.925747+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CherryHQ (CherryHQ/cherry-studio) 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T16:37:40.925747+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Cherry Studio, a cross-platform desktop AI chat application, registered three new Volcano Engine models — doubao-seed-2-1-pro-260628, doubao-seed-2-1-turbo-260628, and doubao-seed-evolving — in its provider registry wit…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Cherry Studio v2 users can now select Seed 2.1 Pro from the doubao provider model list and adjust thinking levels via the Thinking button. The PR implements full VolcEngine API parameter compatibility including thinking…

模型作用:Seed-2.1-Pro-Preview 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Seed 2.1 Pro is integrated as a deep-thinking model with configurable reasoning depth in Cherry Studio. The model's explicit thinking-mode API (thinking.type enabled/disabled) enables users to toggle between fast respon…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Cherry Studio is a third-party open-source project. Multiple competing PRs were submitted simultaneously (#16458, #16463, #16464), suggesting strong community interest in the model.

原始记录:Cherry Studio desktop AI chat client adds Doubao-Seed-2.1 with thinking-level control

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AIYLT 使用 o3-pro 处理金融和商业分析

AIYLT · o3-pro

A
厂商:OpenAI 模型:o3-pro 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AIYLT 公开的金融与商业分析案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:AIYLT 使用 OpenAI o3-2025-04-16 模型(即 o3-pro)构建了一套完整的美股日内交易智能策略系统。系统包含 11 个专业策略模块(Oracle、Helios、Chronos、TerraFilter、Minerva、AlphaForge、Fenrir、Hermes、Aegis、Cerberus、EchoLog),按优先级协同工作,集成了 Polygon.io Level-2 实时数据,目标命中率 ≥80%。

公开产物:系统实现了完整的从数据获取到策略决策到 GitHub 自动推送的全流程自动化。11 个模块分别负责市场扫描、策略评分、环境过滤、行为验证、基本面分析、暗池资金检测、新闻事件分析、风险控制、异常防护和策略回测。代码仓库包含完整的 Python 实现,包括主执行程序 run_o3_pro.py、Polygon API 客户端、GitHub 推送模块和所有策略模块。

模型作用:o3-pro 在此系统中作为核心推理引擎,负责所有策略模块的决策推理。系统设计严格要求使用 o3-pro 模型(绝不降级),利用其深度推理能力进行多维度股票分析和交易决策,包括整合 11 个模块的分析结果进行最终的 Oracle 决策。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该案例来自 GitHub 公开仓库,代码完整可验证。仓库 stars 为 0,可能是较新的个人项目。o3-pro 在此作为 API 调用模型使用(通过 openai Python SDK),非 ChatGPT 界面使用。系统声称目标命中率 ≥80%,但无独立回测结果验证。模型版本绑定为 o3-2025-04-16。

原始记录:AIYLT 构建基于 o3-pro 的美股智能交易策略系统

已有真实案例 金融与商业分析公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Firecrawl (mendableai) 使用 Grok 4 处理知识检索和问答

Firecrawl (mendableai) · Grok 4

A
厂商:xAI / Grok 模型:Grok 4 来源平台:github 最后复核:2026-06-26T14:14:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Firecrawl (mendableai) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T14:14:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Firecrawl built Grok 4 Fire Enrich, an open-source multi-agent data enrichment tool that transforms a list of emails into rich datasets with company profiles, funding data, tech stacks, and executive information. The sy…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working open-source Next.js 15 web application deployed on Vercel that accepts CSV uploads of emails and returns enriched company data including industry classification, funding stage, employee count, headquarters, te…

模型作用:Grok 4 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Grok 4 serves as the AI model powering the agent base execution layer — orchestrating the multi-agent pipeline, running parallel searches via Firecrawl API, and executing the sequential phased extraction strategy. The R…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 51 stars, created by Firecrawl team (mendableai). MIT licensed. Project is actively maintained with deployable Vercel template.

原始记录:Firecrawl Grok 4 Fire Enrich — AI-Powered Email-to-Company Data Enrichment

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

André (buckster123) 使用 Grok 4 处理文档理解和结构化处理

André (buckster123) · Grok 4

A
厂商:xAI / Grok 模型:Grok 4 来源平台:github 最后复核:2026-06-26T14:14:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

André (buckster123) 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T14:14:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ApexOrchestrator is a Streamlit-powered multi-agent AI orchestration framework that uses xAI's Grok models (Grok-4 and Grok-4-fast-reasoning) as its inference engine. The system bootstraps specialized agents (ApexCoder …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production-deployable Streamlit web application with selectable Grok variants (Grok-4-fast-reasoning, Grok-4) via UI, featuring sandboxed code execution, multi-agent dialectical reasoning with Planner/Critic/Executor …

模型作用:Grok 4 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Grok-4 and Grok-4-fast-reasoning serve as the core inference engines for all agent bootstrapping, tool-augmented interaction, and dialectical reasoning. The framework uses xAI's OpenAI-compatible API and selects Grok va…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:14 stars, individual developer project. Uses xAI API directly. Streamlit-based UI. Active development with Docker support.

原始记录:ApexOrchestrator — Multi-Agent AI Orchestration Framework with Dialectical Swarms on Raspberry Pi

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Omer Berat Sezer 使用 Llama 3.1 405B 处理研究分析和报告生成

Omer Berat Sezer · Llama 3.1 405B

A
厂商:Meta / Llama 模型:Llama 3.1 405B 来源平台:github 最后复核:2026-06-26T22:15:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Omer Berat Sezer 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T22:15:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Detect whether input text is AI-generated or human-written, providing an AI generation score, detailed analysis of detection patterns, and explanation of why text appears AI-generated — using Llama 3.1 405B via AWS Bedr…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Streamlit-based tool that returns an AI generation score (0–100%), textual evaluation of AI vs human authorship, and analysis of detection patterns (repetitive phrasing, lack of personal voice, etc.). Published with b…

模型作用:Llama 3.1 405B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Llama 3.1 405B (accessed via AWS Bedrock) performs the actual content analysis and scoring. Its large parameter count enables nuanced pattern detection that smaller models miss — the author specifically chose the 405B v…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Requires a paid AWS Bedrock account with access to the Llama 3.1 405B model (not free tier). Tool is a sample/demo rather than production SaaS.

原始记录:AI Content Detector — AWS Bedrock + Llama 3.1 405B for AI-generated text detection

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Treasure M 使用 Llama 3.1 405B 处理文档理解和结构化处理

Treasure M · Llama 3.1 405B

A
厂商:Meta / Llama 模型:Llama 3.1 405B 来源平台:github 最后复核:2026-06-26T22:15:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Treasure M 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T22:15:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Generate educational materials — lesson plans, worksheets, study guides, assessments, and activities — customized by subject, topic, and grade level (targeting South African curriculum).

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A live browser-based application (custom-content-generator.onrender.com) that generates grade-appropriate educational content with 5+ prompt templates. Outputs can be downloaded as .txt, PDF, or Word .docx files.

模型作用:Llama 3.1 405B 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:meta-llama/llama-3.1-405b-instruct powers all content generation through OpenRouter.ai's API. The 405B model's instruction-following precision is essential for adhering to grade-level constraints and curriculum-specific…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Client-side only — API key must be provided by user. Demo hosted on Render free tier may have cold-start delays. Single-star repo, small project.

原始记录:Custom Content Generator — Educational content creation with Llama 3.1 405B via OpenRouter

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

esparpol4-eng (独立开发者/运营者) 使用 MiMo-V2.5-Pro 处理软件工程任务执行

esparpol4-eng (独立开发者/运营者) · MiMo-V2.5-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2.5-Pro 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

esparpol4-eng (独立开发者/运营者) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发者部署了 Hermes 9Router 多层级 AI 路由网关于腾讯云 VPS,将小米 MiMo V2.5 Pro 设为 Tier 1 首选模型,为 Telegram AI Agent(Hermes)及 Claude Code、Cursor 等客户端提供统一的 OpenAI 兼容接口。实现 7 级顺序回退策略,MiMo 挂时自动切换到后续 Tier。

公开产物:连续运行 11+ 天,累计处理 8,774 个请求、1.92 亿 token。日峰值 2,256 请求 / 54.9M token。MiMo V2.5 Pro 占成功请求的 57.3%(728 个成功请求中 417 个),平均端到端延迟 ~1.7 秒。uptime >99%,systemd 自动重启。估算 30 天费用 ~$814(若全用付费模型)。

模型作用:MiMo V2.5 Pro 作为生产环境首选推理模型,处理多语言(印尼语/中文/英语)Telegram 对话、代码编辑、工具调用等任务。其免费额度(100T plan)和低延迟使其成为成本效益最优的 Tier 1 选择。CJK 原生支持确保非英语提示无质量退化。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:第三方个人部署的生产系统,非小米官方客户案例。数据来自开发者自述的 SQLite 数据库快照,未经独立审计。MiMo API key 为 Token Plan 账户。

原始记录:Hermes 9Router: 腾讯云生产环境 AI 网关以 MiMo V2.5 Pro 为首选 Tier 1 模型

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

NickYoung618 (开发者/研究者) 使用 MiMo-V2.5-Pro 处理代码审查和测试生成

NickYoung618 (开发者/研究者) · MiMo-V2.5-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2.5-Pro 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

NickYoung618 (开发者/研究者) 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发者 NickYoung618 构建了面向湖湘文化遗产的智能问答平台,采用 LangGraph 状态机编排的多智能体架构,融合 Neo4j 图谱检索、ChromaDB 向量语义检索和 Tavily 网络搜索三路并行检索。MiMo-V2.5-Pro 作为核心 LLM,负责实体抽取、意图识别、图谱查询生成和流式回答生成。

公开产物:完整的全栈项目:Python FastAPI 后端(SSE 流式输出)+ Next.js 16 前端 + Neo4j 图数据库 + ChromaDB 向量库。支持 152+ 种子实体的湖湘文化知识图谱、多智能体角色(协调员/研究员/遗产专家/分析师/讲述者)、ECharts 交互式图谱可视化。已部署并可本地运行。

模型作用:MiMo V2.5 Pro 承担所有 LLM 推理工作:(1) 从用户提问中抽取实体并识别意图,生成精准 Cypher 查询;(2) 融合三路检索结果后进行流式生成回答;(3) 在多智能体角色中作为协调员进行路由分发。其长上下文能力和中文理解能力使其适合处理大规模文化遗产知识。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:学术/个人项目,尚未公开部署实例。依赖 MiMo API(通过 token-plan-cn 端点)。Neo4j 和 ChromaDB 数据量有限(12 个清洗后百科词条)。

原始记录:湖湘文化遗产数字化交互系统: 基于 MiMo-V2.5-Pro 的 GraphRAG 多智能体知识问答平台

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

moriarthur (Doclos) 使用 GLM-4.7 处理知识检索和问答

moriarthur (Doclos) · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:github 最后复核:2026-06-27T00:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

moriarthur (Doclos) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T00:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Doclos is a SaaS platform that automates business document processing for German Mittelstand (small and medium enterprises). Users upload PDF invoices (Rechnungen), contracts (Verträge), offers (Angebote), and delivery …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production SaaS platform that automatically processes German business documents — invoices, contracts, offers, delivery notes — extracting structured, searchable data from PDFs into a PostgreSQL database with JWT auth…

模型作用:GLM-4.7 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7-Flash serves as the core AI engine in Doclos's extraction pipeline, responsible for understanding document layout and semantics, extracting key fields (dates, amounts, parties, line items) with per-field confide…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo is open-source on GitHub with clear README and architecture docs. Model usage is explicitly documented in the tech stack table. No official Zhipu customer story — this is a third-party developer's production build.

原始记录:Doclos: AI-powered document automation SaaS extracting structured data from German business PDFs using GLM-4.7-Flash

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

graphanov / john-lomein (automated … 使用 GLM-4.7 处理软件工程任务执行

graphanov / john-lomein (automated maintainer) · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

graphanov / john-lomein (automated maintainer) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LazyGLM is a standalone terminal coding agent CLI for GLM models, providing a Claude Code-style interactive REPL for AI-assisted software development. It routes tasks by type to different GLM model tiers: GLM-5.2 for ha…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An npm-installable CLI tool (lazyglm) that provides an interactive agentic coding shell with GLM model routing, session persistence, thinking-token replay across turns, and headless subprocess mode for CI/CD integration…

模型作用:GLM-4.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7 is the designated model for routine coding tasks, completion verification, and quick-edit turns within LazyGLM's routing system. It handles the 'verifier' and 'quick' roles, providing efficient, cost-effective c…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:New project (0 stars as of collection). Automated maintainer bot. No formal license stated in README.

原始记录:LazyGLM: GLM-Native Coding Agent CLI with Model Routing to GLM-4.7

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kittycheeseburger 使用 GLM-4.7-Flash 处理研究分析和报告生成

kittycheeseburger · GLM-4.7-Flash

A
厂商:Z AI / GLM 模型:GLM-4.7-Flash 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kittycheeseburger 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A front-end/back-end separated emotion analysis chatbot web application that uses GLM-4.7-Flash as the default conversation engine. The system combines GLM-4.7-Flash for dialogue generation with a local Chinese RoBERTa …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A full-stack web app (FastAPI backend + Vite/React frontend) that provides real-time emotion-aware chat. GLM-4.7-Flash generates conversational responses, while a local Transformer model classifies user emotions (return…

模型作用:GLM-4.7-Flash 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7-Flash is the primary conversation generation model, producing natural dialogue responses in the emotion analysis chatbot. It is called via the Zhipu Open Platform API (open.bigmodel.cn) and configured as the def…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Student/personal project. No deployment URL provided. API key required for GLM backend.

原始记录:Emotion Chatbot: Full-Stack Emotion Analysis Chatbot Powered by GLM-4.7-Flash

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

tljcpa 使用 GLM-4.7-Flash 处理研究分析和报告生成

tljcpa · GLM-4.7-Flash

A
厂商:Z AI / GLM 模型:GLM-4.7-Flash 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tljcpa 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:An LLM-powered CLI tool for automated academic paper polishing and Word document formatting. The tool uses GLM-4.7-Flash to perform academic-grade content rewriting, grammar correction, and structural recognition on rou…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Python CLI tool (main.py) that takes a rough thesis draft (input.txt), calls GLM-4.7-Flash for academic content polishing via the Zhipu API, applies chunked processing for long documents, and outputs a properly format…

模型作用:GLM-4.7-Flash 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7-Flash is the default and recommended backend for the paper polishing task, chosen for its domestic accessibility in China, free tier availability (200K context), and strong Chinese text understanding capability.…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Student project for undergraduate thesis formatting. Free tier rate-limited (~1 QPS). Requires Zhipu API key with real-name verification.

原始记录:Paper Polisher MVP: Academic Thesis Polishing and Auto-Formatting Tool Using GLM-4.7-Flash

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ByteDance 使用 DeepSeek V3.2 处理研究分析和报告生成

ByteDance · DeepSeek V3.2

A
厂商:DeepSeek 模型:DeepSeek V3.2 来源平台:github 最后复核:2026-06-26T22:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ByteDance 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T22:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ByteDance open-sourced DeerFlow, a long-horizon SuperAgent harness that researches, codes, and creates using sandboxes and memory. The official README explicitly recommends DeepSeek V3.2 as one of the top models to run …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:DeerFlow is deployed as an autonomous agent framework with 74k+ GitHub stars. When configured with DeepSeek V3.2, the model handles multi-step research, code generation, and content creation tasks within the agent pipel…

模型作用:DeepSeek V3.2 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.2 serves as one of the three officially recommended backbone LLMs for DeerFlow. Its thinking-in-tool-use capability and strong reasoning make it suitable for the framework's long-horizon tasks requiring mult…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:DeerFlow supports multiple models; V3.2 is recommended but not the only option. The README states 'We strongly recommend using Doubao-Seed-2.0-Code, DeepSeek v3.2 and Kimi 2.5 to run DeerFlow'.

原始记录:ByteDance DeerFlow: Long-Horizon SuperAgent Framework Powered by DeepSeek V3.2

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Zeronuru404 使用 MiMo-V2.5-Pro 处理软件工程任务执行

Zeronuru404 · MiMo-V2.5-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2.5-Pro 来源平台:github 最后复核:2026-06-26T22:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Zeronuru404 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Zeronuru404 构建了 NusaShield AI——一个面向印尼市场的多智能体欺诈检测平台,专门针对 QRIS(印尼快速支付)、钓鱼攻击和发票欺诈三大场景。核心推理引擎使用 mimo-v2.5-pro 长链推理能力,三个高复杂度 Agent(QRIS Scanner 600K tokens/次、Phishing Detector 400K tokens/次、Invoice Fraud Analyzer 700K tokens/次)均绑定 mimo-v2.5-pro 模型,每天处理约 43.9M tokens。系统参考印尼央行(BI)法规、OJK 指令和 BSSN 安全公告进行欺诈模式识别。

公开产物:六智能体协作的端到端欺诈检测系统:QRIS Scanner 检测 QRIS swap 和伪造商户码、Phishing Detector 分析 SMS/WhatsApp/Email 钓鱼链接、Invoice Fraud Analyzer 解析发票 PDF 并交叉验证 NPWP/银行账号、WhatsApp Scam Agent 分类 WhatsApp Business 诈骗、Threat Intel Agent 跨 BI/OJK/BSSN 情报源交叉验证、Report Generator 输出印尼语执行摘要。技术栈包括 Hermes Agent 框架、PostgreSQL、Redis、Docker+Kubernetes。

模型作用:MiMo-V2.5-Pro 是系统的核心推理引擎,承担三个最高复杂度 Agent(QRIS Scanner、Phishing Detector、Invoice Fraud Analyzer)的长链推理任务。每个 Agent 单次运行消耗 400K-700K tokens,需要模型具备跨语言(印尼语)多步推理、法律条文理解和欺诈模式识别能力。MiMo-V2.5-Pro 的 1M 上下文窗口和长链推理能力是选择该模型的关键原因。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目 README 标注 'Development in progress — seeking MiMo API credits for production scale',尚未进入生产部署阶段。项目为开发中状态,有完整 README 架构设计和 token 消耗模型,但缺少实际生产运行截图或用户反馈。

原始记录:NusaShield AI: 印尼金融科技多智能体欺诈检测系统,以 MiMo-V2.5-Pro 驱动 QRIS/钓鱼/发票欺诈推理

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ricoredfinger-png 使用 MiMo-V2.5-Pro 处理智能体流程编排

ricoredfinger-png · MiMo-V2.5-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2.5-Pro 来源平台:github 最后复核:2026-06-26T22:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ricoredfinger-png 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T22:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ricoredfinger-png 构建了 Praxis Forge——一个自主多智能体代码库现代化和技术债务清理系统。系统包含 7 个专业 Agent(Scout 扫描、Auditor AST 分析、Mason 重构、Smith 测试合成、Scribe 文档、Watch CI 门禁、Forge PR 作者),每天消耗 19.8B tokens,通过 4 轮长链推理(4-pass long-chain reasoning)做出每个重构决策。MiMo-V2.5-Pro 作为 Forge Orchestrator 负责全局协调和记忆管理。

公开产物:7 专业 Agent 组成的自主代码维护流水线:Scout Agent 扫描仓库技术债务、Auditor Agent 通过 AST 分析代码异味、Mason Agent 执行重构、Smith Agent 合成测试、Scribe Agent 生成文档和注释、Watch Agent 管理 CI 门禁、Forge Agent 生成 PR。整个系统每天消耗约 19.8B tokens,持续监控代码库并自动提交清洁 PR。支持 Docker 部署,提供 69 场景工具评估,质量评分 97.8/100。

模型作用:MiMo-V2.5-Pro 作为 Forge Orchestrator 承担全局协调和记忆管理职责,是整个 7-Agent 系统的中枢。其长链推理能力用于 4-pass 决策链——先理解遗留代码存在的原因,再制定重构策略,确保不是简单的 find-and-replace。系统标注 'Powered by Xiaomi MiMo V2.5' 作为核心技术依赖。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目 README 标注每日 19.8B tokens 消耗量,但未明确是否已在生产环境实际部署运行。提供 proof bundle 截图和 69 场景工具评估结果,有 Docker 部署支持,但缺少第三方验证或用户反馈。

原始记录:Praxis Forge: 7 智能体自动代码库现代化系统,MiMo-V2.5-Pro 作为核心编排器

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ph419 使用 MiMo-V2.5-Pro 处理软件工程任务执行

ph419 · MiMo-V2.5-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2.5-Pro 来源平台:github 最后复核:2026-06-26T22:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ph419 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ph419 开发了 Tackle Harness——一个基于插件的 AI Agent 工作流框架,为 Claude Code 提供任务管理、工作流编排、角色管理等能力。框架支持多 provider 切换,MiMo-V2.5-Pro 是其中的核心执行模型之一。框架通过 provider-resolver 自动探测 mimo 模型并启用纯透传模式(无额度约束),支持多 loop 并行和全局协调。用户通过 `--settings=~/.claude/mimo-v2.5-pro.json` 配置使用 MiMo-V2.5-Pro 执行复杂的多阶段任务(规划→审核→执行→验证→交付)。

公开产物:五阶段端到端任务交付流水线:P0 规划(task-creator / split-wp 拆分工作包)→ P1 人工审核(human-checkpoint)→ P2 执行(agent-dispatcher 调度多个 Agent 并行工作)→ P3 验证 → P4 交付。支持多 loop 并行运行、全局协调器(loop-server 守护进程)聚合多 loop 视图和按 provider 分桶统筹额度池。框架提供经验沉淀机制,每次任务完成后自动提炼经验教训供后续任务参考。

模型作用:MiMo-V2.5-Pro 是 Tackle Harness 的核心执行模型之一。框架的 provider-resolver 自动识别 mimo 模型并启用纯透传模式,支持用户通过 settings JSON 配置 MiMo-V2.5-Pro 的端点和认证。在多 provider 架构中,MiMo-V2.5-Pro 被作为高性价比的长链推理选择,与 GLM、DeepSeek 等并列支持。框架设计充分利用了 MiMo-V2.5-Pro 的 1M 上下文窗口进行复杂任务规划。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目 26 stars,规模较小,但活跃更新(最后更新 2026-06-26)。README 文档详实,有明确的 MiMo-V2.5-Pro 集成文档和配置示例。作为框架工具而非终端应用,更多是展示 MiMo-V2.5-Pro 作为 provider 被集成的生态价值。

原始记录:Tackle Harness: AI Agent 工作流框架将 MiMo-V2.5-Pro 作为核心执行模型用于任务编排与自动交付

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

HOWILLMAKEIT 使用 DeepSeek V3.2 处理代码审查和测试生成

HOWILLMAKEIT · DeepSeek V3.2

A
厂商:DeepSeek 模型:DeepSeek V3.2 来源平台:github 最后复核:2026-06-26T15:04:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

HOWILLMAKEIT 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T15:04:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:轻量通用中文 RAG 助手,支持多知识库管理:后端 FastAPI + LlamaIndex + FAISS,嵌入使用 Qwen text-embedding-v4,生成使用 DeepSeek-V3.2-Exp 的非思考模式。前端 Vite + React 提供知识库管理和对话问答页面。

公开产物:完整可用的 RAG 产品,含网页端和 Electron 桌面端。桌面端已发布 Windows 安装包,支持本地部署、知识库切换、AES 加密配置管理。

模型作用:DeepSeek-V3.2-Exp 非思考模式作为核心生成模型,负责基于检索到的上下文生成高质量中文回答。选择非思考模式以实现低延迟的实时问答体验。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人开发者项目,Stars 9。使用 V3.2-Exp 而非 V3.2 正式版。

原始记录:EasyRAG: 通用中文 RAG 助手使用 DeepSeek-V3.2-Exp 非思考模式生成答案

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chuancyzhang 使用 DeepSeek V3.2 处理文档理解和结构化处理

chuancyzhang · DeepSeek V3.2

A
厂商:DeepSeek 模型:DeepSeek V3.2 来源平台:github 最后复核:2026-06-26T15:04:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chuancyzhang 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T15:04:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:基于 DeepSeek-V3.2 的 Windows 桌面 Agent 工作台,集成对话、项目文件操作、技能扩展、自动化模板、定时任务和长期记忆。支持文件读写、工具调用、多步骤任务完成,以及 Markdown/HTML/PDF/DOCX/PPTX/XLSX 内容预览。

公开产物:成熟的桌面应用(v4.9.4),PySide6 GUI,支持纯聊天和项目绑定两种模式。包含内置技能、可选插件、用户技能和 MCP 工具扩展体系。

模型作用:DeepSeek-V3.2 的推理能力和工具调用能力是整个 Agent 框架的核心。模型负责理解用户指令、规划多步骤任务、调用文件系统和外部工具、以及生成结构化的操作结果。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人开发者项目,Stars 9。项目声明非 DeepSeek 官方产品。

原始记录:DeepSeek Cowork: Windows 桌面 Agent 工作台使用 V3.2 推理+工具调用

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Qcunyu 使用 DeepSeek V3.2 处理知识检索和问答

Qcunyu · DeepSeek V3.2

A
厂商:DeepSeek 模型:DeepSeek V3.2 来源平台:github 最后复核:2026-06-26T15:04:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Qcunyu 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T15:04:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:基于 71k 字原始语料、15 部核心著作集、10+ 场访谈与百余封书信,通过多轮蒸馏将鲁迅的二难推理、归谬法和国民性剖析逻辑植入 LLM 推理链。集成"女娲"Skill 体系,采用 DeepSeek-V3.2 Reasoner 建模。

公开产物:可用的认知框架 Skill,提供 5 个核心心智模型、4 条决策启发式和排他性表达 DNA。支持社会文化现象的深度剖析,内置批判逻辑工具。

模型作用:DeepSeek-V3.2 Reasoner 作为基础模型,利用其强推理能力进行知识蒸馏。Reasoner 模式的长链推理特性使得鲁迅式的多层论证和批判逻辑能够被有效复刻和迁移。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人项目,Stars 6。属于知识蒸馏/认知框架复刻,非直接的生产应用。

原始记录:Deep-LuXun-skill: 基于 V3.2 Reasoner 深度蒸馏的鲁迅认知框架复刻

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DrowskoytayhulGuider 使用 DeepSeek V3.2 处理真实任务执行

DrowskoytayhulGuider · DeepSeek V3.2

A
厂商:DeepSeek 模型:DeepSeek V3.2 来源平台:github 最后复核:2026-06-26T15:04:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DrowskoytayhulGuider 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T15:04:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:接入 DeepSeek-V3.2 API 注入丁真角色提示词,结合 MockingBird 训练的 TTS 语音模型和 Vosk 语音识别,实现带有语音交互的整活聊天机器人。支持 GUI 和命令行两种交互模式。

公开产物:可用的聊天机器人,含 GUI 界面和命令行界面。支持语音识别输入、角色扮演对话、丁真语音合成输出。提供雪豹闭嘴/纯纯出声等趣味交互功能。

模型作用:DeepSeek-V3.2 API 作为对话生成的核心引擎,负责根据角色提示词生成符合丁真人设的回复内容。V3.2 的指令遵循能力使得角色一致性在多轮对话中得以保持。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:趣味/整活项目,Stars 12。角色扮演场景,非严肃生产应用。依赖 DeepSeek 官方 API 的可用性。

原始记录:电子丁真: DeepSeek-V3.2 API 驱动的角色扮演聊天机器人 + MockingBird TTS

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

karanb192 (independent developer) 使用 o3-pro 处理真实任务执行

karanb192 (independent developer) · o3-pro

A
厂商:OpenAI 模型:o3-pro 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

karanb192 (independent developer) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer karanb192 built and published an Open WebUI function/plugin that integrates OpenAI o3-pro reasoning models into the Open Web UI platform, providing comprehensive per-message token usage tracking, cumulative co…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Published open-source plugin (11 stars, 8 forks) on GitHub and the Open Web UI marketplace at openwebui.com/f/karanb192/o3pro_o1pro_support. The plugin enables o3-pro usage in Open Web UI with real-time cost tracking sh…

模型作用:o3-pro 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o3-pro is the primary reasoning model integrated by the plugin. The plugin was specifically designed around o3-pro's unique characteristics: its Responses API requirement (not standard chat/completions), its $20/$80 per…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a tool/integration case rather than a direct task-solving case. The developer built the plugin to enable o3-pro access rather than using o3-pro to solve a specific domain problem. However, the plugin is actively…

原始记录:Developer builds Open WebUI plugin for o3-pro reasoning model access with cost tracking

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

serenemm (Roblox developer) 使用 MiMo-V2-Flash (Feb 2026) 处理可玩交互原型构建

serenemm (Roblox developer) · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

serenemm (Roblox developer) 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Deploy an LLM-powered NPC named 'Sam' inside a Roblox game that holds real-time natural language conversations with players via in-game chat bubbles, using MiMo-V2-Flash as the inference backend served through OpenRoute…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working Roblox game with an AI NPC ('Sam') that receives player chat messages through HttpService, sends them to a Railway-hosted backend, queries xiaomi/mimo-v2-flash:free via OpenRouter, and returns contextual conve…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash provides the core conversational intelligence: it processes natural language player input and generates contextually appropriate NPC responses. The model was chosen for its free tier availability on OpenRo…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:README architecture section mentions 'deepseek LLM model' in one place, likely a copy-paste error from a template; the explicit model declaration throughout the README and setup instructions consistently states xiaomi/m…

原始记录:Project Sam: AI-Powered Roblox NPC Conversations via MiMo-V2-Flash

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

herointene (Discord bot developer) 使用 MiMo-V2-Flash (Feb 2026) 处理翻译和本地化处理

herointene (Discord bot developer) · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

herointene (Discord bot developer) 公开的翻译与本地化案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕翻译和本地化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Discord bot that provides context-aware multilingual translation between Chinese, Japanese, and English, with automatic AI task recognition (summarization, email writing, etc.), using MiMo-V2-Flash as the sole L…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production-ready Discord bot deployed via Docker Compose with: (1) Flag emoji reaction triggers (🇨🇳/🇯🇵/🇬🇧) for forced target-language translation, (2) thread-based interaction to keep channels clean, (3) smart c…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担翻译和本地化处理相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash is the core engine powering all AI capabilities: multilingual translation, context understanding from recent message history, intent classification (translation vs. task fulfillment), and generation of tas…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:No stars or external adoption metrics visible yet. The bot is functional but appears to be an individual developer project without documented production deployment scale.

原始记录:Discord AI Translator: Context-Aware Multilingual Bot Powered by MiMo-V2-Flash

已有真实案例 翻译与本地化公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AndreaPanzeriDev (game developer) 使用 MiMo-V2-Flash (Feb 2026) 处理可玩交互原型构建

AndreaPanzeriDev (game developer) · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AndreaPanzeriDev (game developer) 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Use MiMo-V2-Flash to generate 100% of the source code for a fully playable 2D tower defense game, including 4 tower types, 3 enemy types, wave management, economy system, and Docker deployment configuration.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete 2D tower defense game with: 4 tower types (Basic, Sniper, Splash, Slow), 3 enemy types (Normal, Fast, Tank), increasing-difficulty wave system with UI, economy system (money from kills, tower costs/upgrades/s…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash generated the entire codebase including game logic, rendering, UI, enemy AI, tower mechanics, wave spawning, economy balancing, and Docker configuration. The developer used a single prompt ('realize a 2d g…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Zero stars; appears to be a personal demo project. The '100% coded' claim is asserted by the developer but the repo contains only a Dockerfile and game code without explicit generation logs.

原始记录:Tower Defense Game 100% AI-Coded Using MiMo-V2-Flash

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

blessblissmari (academic tool devel… 使用 MiMo-V2-Flash (Feb 2026) 处理代码审查和测试生成

blessblissmari (academic tool developer) · MiMo-V2-Flash (Feb 2026)

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Flash (Feb 2026) 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

blessblissmari (academic tool developer) 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a 5-pass cascade translation pipeline that converts English academic PDF/DOCX/PPTX documents into Russian DOCX/PPTX in MIET university template format, where MiMo-V2-Flash serves as the cost-efficient OCR comparis…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A web application (v1 live at blessblissmari.github.io/miet-translator-pro/) with a v2 cascade pipeline: Pass 1 (mimo-v2.5, global OCR layout map), Pass 2 (mimo-v2-omni, per-page OCR), Pass 3 (MiMo-V2-Flash, diff compar…

模型作用:MiMo-V2-Flash (Feb 2026) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash is used as the 'CHEAP' tier in a sophisticated multi-model pipeline, handling two critical roles: (1) Pass 3 'SVERKA' — diff-comparing OCR output against pdf.js text extraction to catch translation artifac…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Live demo at blessblissmari.github.io/miet-translator-pro/ returned empty response during verification (may be temporarily down or require JavaScript). GitHub repo README and source code structure are fully verifiable. …

原始记录:MIET Translator Pro v2: Academic Document Translation Pipeline Using MiMo-V2-Flash as QA Layer

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OmuNaman (open-source developer) 使用 Gemini 3.1 Pro Preview 处理软件工程任务执行

OmuNaman (open-source developer) · Gemini 3.1 Pro Preview

A
厂商:Google / Gemini 模型:Gemini 3.1 Pro Preview 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OmuNaman (open-source developer) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Open-source terminal tool that watches a specified subreddit every 30 minutes, scrapes recent posts (titles, bodies, full comment threads with nested replies), downloads attached meme images, and sends everything as a b…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Produces intent-driven intelligence briefs in Markdown and JSON format, containing sentiment analysis, recurring themes, identified brands/people/institutions, community perception narratives, and evidence citations (Po…

模型作用:Gemini 3.1 Pro Preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 3.1 Pro Preview is the sole analysis engine. Its 1M-token context window enables batch-packing hundreds of Reddit posts with full comment threads into single analysis passes. Native multimodal capability processe…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Indie developer project, not enterprise production deployment. No stars visible in search results yet. However, the code is well-structured, functional, and the .env.example explicitly configures gemini-3.1-pro-preview …

原始记录:Reddit Community Intelligence Scanner: Multimodal Sentiment Analysis with Gemini 3.1 Pro Preview

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

86thAuspiciousVerse 使用 DeepSeek V4 Pro (Max) 处理软件工程任务执行

86thAuspiciousVerse · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

86thAuspiciousVerse 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer used DeepSeek V4 Pro Max to port the fluid_vb embedded firmware project from CH32V203 RISC-V platform to RP2040-Zero (RP2040) board. The repository was explicitly created to test V4 Pro Max's coding capability…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Complete C++ codebase for RP2040-Zero implementing fluid simulation firmware originally written for CH32V203. Repository contains working ported code with hardware abstraction for RP2040 GPIO/peripheral differences.

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro Max was used as the primary coding assistant for the entire firmware porting task — generating C++ code that adapts hardware registers, peripheral initialization, and control logic from one RISC-V MCU to…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Single-commit repo (2026-05-04), 0 stars. Small personal project; code quality and correctness not independently verified. The repo description explicitly states it was created to test V4 Pro Max coding capability.

原始记录:DeepSeek V4 Pro Max used to port embedded firmware from CH32V203 to RP2040-Zero

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

KYJCASTER 使用 DeepSeek V4 Pro (Max) 处理软件工程任务执行

KYJCASTER · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T22:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

KYJCASTER 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer used DeepSeek V4 Pro Max to build an optimized version of the feedsystem_video_go project — a Go + Vue 3 short video feed system with account management, video upload/publish, likes, comments, follows, feed ra…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Complete production-grade full-stack application with Go backend (GORM, JWT auth, Redis caching, RabbitMQ async workers), Vue 3 frontend, Docker Compose deployment, 100 test users, and comprehensive API documentation co…

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro Max was used as the primary development tool to build/optimize the entire system. Commit history shows rapid development session (5 commits in ~45 minutes on 2026-04-25) covering features from P2 (SSE no…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:0 stars, single development session. README does not explicitly mention DeepSeek usage — only the repository description states 'deepseek v4 pro(max) version'. The original project by LeoninCS/feedsystem_video_go exists…

原始记录:DeepSeek V4 Pro Max used to build a full-stack short video feed system (Go + Vue 3)

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Doriandarko 使用 Kimi K2 Thinking 处理软件工程任务执行

Doriandarko · Kimi K2 Thinking

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 Thinking 来源平台:github 最后复核:2026-06-27T05:49:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Doriandarko 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:49:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build and run a Python autonomous writing agent that accepts a story or book prompt, plans the writing project, creates project folders, writes Markdown files, streams generation, manages long context, and supports reco…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository for Kimi Writing Agent, described as powered by kimi-k2-thinking, with runnable CLI usage examples for creating sci-fi short story collections and recovery workflows; the README documents Moonsh…

模型作用:Kimi K2 Thinking 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2 Thinking is the core agent model: it plans the writing task, produces the prose, decides when to call file-management tools, and summarizes or compresses context so longer writing projects can continue.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is the maintainer's public GitHub README and reachable repository; no independent production metrics are published, but the repo is a concrete runnable artifact rather than a benchmark, tutorial-only page, laun…

原始记录:Doriandarko built an autonomous creative-writing agent with Kimi K2 Thinking

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

HarleyCoops 使用 Kimi K2 Thinking 处理软件工程任务执行

HarleyCoops · Kimi K2 Thinking

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 Thinking 来源平台:github 最后复核:2026-06-27T05:49:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

HarleyCoops 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:49:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a multi-agent pipeline that turns mathematical or physics concepts into Manim animation specifications and rendered explainer animations, including prerequisite exploration, mathematical enrichment, visual planni…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository containing the KimiK2Manim package, pipeline documentation, generated GIF examples such as minimal-surface and Brownian-motion scenes, and Moonshot API configuration/docs links for Kimi K2 thinking mod…

模型作用:Kimi K2 Thinking 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2 Thinking supplies the structured reasoning and tool-calling steps for the agent pipeline, extracting prerequisites, mathematical content, visual specifications, and narrative prompts used to generate Manim scene…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence comes from a public GitHub project README and reachable repository; the project explicitly describes using the Kimi K2 thinking model from Moonshot AI for the artifact, though it is an unofficial user project r…

原始记录:HarleyCoops used Kimi K2 Thinking to generate Manim math and physics explainer animations

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

prnake 使用 Kimi K2 Thinking 处理软件工程任务执行

prnake · Kimi K2 Thinking

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 Thinking 来源平台:github 最后复核:2026-06-27T05:49:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

prnake 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:49:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Implement a runnable agentic search and browsing workflow for Kimi K2 Thinking that can answer complex research questions or generate news reports by repeatedly searching, reading retrieved pages, folding prior tool res…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository with CLI and frontend entry points, sample official trace data, example prompts for daily news reporting and complex multi-hop web questions, and documentation explaining the Kimi K2 Thinking ag…

模型作用:Kimi K2 Thinking 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2 Thinking is used as the reasoning and orchestration model that decides search and browse steps, maintains the tool-use trajectory, and composes the final researched answer while external search/page-fetch tools …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Unofficial personal implementation, not an official Moonshot customer story; nevertheless it is a concrete public runnable artifact tied to Kimi K2 Thinking rather than a benchmark, launch summary, prompt-only page, or …

原始记录:prnake recreated Kimi K2 Thinking agentic search as a runnable deep-research tool

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Workshop Labs (workshop-labs-pbc) 使用 Kimi K2 Thinking 处理真实任务执行

Workshop Labs (workshop-labs-pbc) · Kimi K2 Thinking

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 Thinking 来源平台:github 最后复核:2026-06-27T03:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Workshop Labs (workshop-labs-pbc) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T03:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Workshop Labs 团队在 HuggingFace 上对 Kimi K2 Thinking 进行 LoRA 微调的完整实践。需要解决的核心挑战:Kimi K2 使用量化专家权重(NF4),HuggingFace 原生不支持训练量化专家。团队通过 monkey-patch 方式绕过限制,使用 8×H200 GPU 在自定义 Yoda 风格数据集上训练 40 步,验证损失下降且模型行为发生预期变化。

公开产物:成功在 HuggingFace 框架上对 Kimi K2 Thinking 进行 LoRA 微调。训练在 8×H200 GPU 上完成,40 步训练后损失下降。产出包含完整的训练代码(train.py)、数据生成脚本(make_yoda_dataset.py)、训练日志和损失曲线可视化。技术博客详细记录了遇到的 bug 和解决方案。

模型作用:Kimi K2 Thinking 的开放权重使社区能够进行微调实验。模型的 MoE 架构(384 专家,每 token 选 8 个 + 1 个共享专家)对微调提出了独特挑战,团队需要 monkey-patch HuggingFace 代码以支持量化专家训练。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是微调/训练场景而非推理使用。需要 8×H200 GPU。博客文章 URL 可能有访问限制。

原始记录:Workshop Labs 使用 LoRA 微调 Kimi K2 Thinking 的实践

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Moonshot AI Research Team 使用 Kimi K2.5 处理代码审查和测试生成

Moonshot AI Research Team · Kimi K2.5

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.5 来源平台:official_web 最后复核:2026-06-26T22:55:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Moonshot AI Research Team 公开的代码审查与测试案例,来源为 官方页面,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Evaluate Kimi K2.5 on WorldVQA, a novel benchmark testing atomic visual knowledge of frontier models across 9 task categories (People, Objects, Culture, Geography, Transportation, Brands, Sports, Entertainment, Location…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Kimi K2.5 achieved 46.3% accuracy on WorldVQA Overall Accuracy, ranking second only to Gemini-3-pro (47.4%) and outperforming Claude-opus-4.5 (36.8%), GPT-5.2 (28.0%), Qwen3-VL-235B (23.5%), and all other evaluated mode…

模型作用:Kimi K2.5 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.5 demonstrates state-of-the-art visual knowledge understanding, achieving near-Gemini-3-pro performance on a challenging benchmark that tests real-world visual knowledge rather than OCR or simple recognition. Th…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Published on Moonshot AI's official blog with arxiv paper reference. Benchmark data is from the model provider's own research but with open-source reproducible evaluation.

原始记录:K2.5 Powering WorldVQA Visual Knowledge Benchmark Research

已有真实案例 代码审查与测试官方页面A 类可核验real_case auto_approved 进入模型卡精选

Stray Labs (straylabs-ai) 使用 Kimi K2.5 处理代码审查和测试生成

Stray Labs (straylabs-ai) · Kimi K2.5

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.5 来源平台:github 最后复核:2026-06-27T03:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Stray Labs (straylabs-ai) 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T03:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:使用 Kimi K2.5 模型构建自主 Web 应用渗透测试代理 Deadend CLI。该工具采用反馈驱动迭代架构(ADaPT),在标准工具失败时自动生成自定义 Python 漏洞利用代码,观察响应并迭代优化攻击策略。支持完全本地执行,无数据外泄。

公开产物:Deadend CLI 在 XBOW 104 题验证套件上使用 Kimi K2.5 达到约 80% 的成功率,总 API 成本约 $122。在 SQL 注入(83%)、XSS(91%)、业务逻辑(86%)等类别表现突出,GraphQL 和 SSRF 达到 100%。在盲注 SQL 注入等其他代理得分为 0% 的挑战上成功突破。

模型作用:Kimi K2.5 提供了强大的代码生成和推理能力,使代理能够自主生成自定义漏洞利用代码、分析 Web 应用响应并迭代优化攻击策略。模型的工具调用能力支持与 Playwright、Docker 等沙箱工具的集成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:使用 Kimi K2.5 变体(较新版本)。基准测试结果来自 2026 年 1 月的 XBOW 验证套件。

原始记录:Kimi K2.5 驱动的自主渗透测试代理

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

sangjiexun 使用 Kimi K2.5 处理研究分析和报告生成

sangjiexun · Kimi K2.5

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.5 来源平台:github 最后复核:2026-06-26T23:14:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

sangjiexun 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T23:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer sangjiexun built ClawCoder, a pip-installable interactive CLI coding assistant powered by Kimi K2.5 through the Moonshot AI API (api.kimi.com/coding, model: k2p5). The tool provides AI-powered conversation, fi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully functional CLI tool published on PyPI (pip install clawcoder) with features including REPL mode, directory-aware file operations, AI-driven code editing, binary disassembly, and persistent chat history. The tool…

模型作用:Kimi K2.5 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.5 (via model identifier k2p5) is the sole AI backend powering all intelligent features: natural language conversation, code generation and editing, file analysis, and multi-modal understanding. The 256K context …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small open-source project (7 stars). Developer is an individual contributor. Tool uses Kimi K2.5's proprietary API endpoint.

原始记录:sangjiexun builds ClawCoder CLI coding assistant powered by Kimi K2.5 via LangGraph

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

鲲鹏Talk (Bilibili creator, mid:18052… 使用 Kimi K2.5 处理软件工程任务执行

鲲鹏Talk (Bilibili creator, mid:18052876) · Kimi K2.5

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.5 来源平台:bilibili 最后复核:2026-06-26T23:14:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

鲲鹏Talk (Bilibili creator, mid:18052876) 公开的代码代理与软件工程案例,来源为 bilibili,复核于 2026-06-26T23:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Bilibili tech creator 鲲鹏Talk demonstrated end-to-end development of a fund company portal website using Kimi Code paired with Kimi K2.5 model. The video shows the complete development workflow as a practical replacement for Claude Code, covering frontend layout, backend logic, and full-stack integration. The creator explicitly positioned Kimi K2.5 as a viable alternative to Claude Code for real software development projects.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Successfully built and demonstrated a complete fund company portal website with Kimi Code + K2.5. The video received 64,063 views and 306 likes, indicating the demo was well-received by the developer community. Publishe…

模型作用:Kimi K2.5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.5 serves as the core coding intelligence behind Kimi Code, handling all code generation, debugging, and project scaffolding tasks throughout the website development process. The model's agentic coding capabiliti…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Video demonstration by a single creator. The website is demonstrated in the video but the actual deployed site URL is not provided. Bilibili video is the primary evidence.

原始记录:鲲鹏Talk develops fund company portal website using Kimi Code + K2.5 as Claude Code alternative

已有真实案例 代码代理与软件工程bilibiliA 类可核验real_case auto_approved 进入模型卡精选

阿布AI进化论 (Bilibili creator, mid:4999… 使用 Kimi K2.5 处理多模态内容处理

阿布AI进化论 (Bilibili creator, mid:499994152) · Kimi K2.5

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.5 来源平台:bilibili 最后复核:2026-06-26T23:14:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

阿布AI进化论 (Bilibili creator, mid:499994152) 公开的多模态生成与理解案例,来源为 bilibili,复核于 2026-06-26T23:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Bilibili creator 阿布AI进化论 built and deployed a production AI assistant system by integrating OpenClaw with Kimi K2.5 as the underlying model and Feishu (Lark) bot as the communication interface, creating a 7x24 always-on AI assistant for work. The tutorial demonstrates the complete Windows-based setup from OpenClaw installation through Kimi K2.5 model configuration to Feishu bot webhook deployment.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully functional 7x24 enterprise AI assistant deployed via Feishu bot, powered by Kimi K2.5 through OpenClaw. The video received 25,575 views and 384 likes. Published on 2026-02-11, indicating early-adopter deployment…

模型作用:Kimi K2.5 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.5 is the core reasoning model powering the AI assistant's conversational intelligence, task execution, and multi-turn dialogue capabilities through the OpenClaw framework. The model handles natural language unde…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Tutorial-format video. While it demonstrates real deployment, the primary artifact is a setup guide rather than a production usage report. Published by a single content creator.

原始记录:阿布AI进化论 deploys OpenClaw + Kimi K2.5 + Feishu bot as 7x24 enterprise AI assistant

已有真实案例 多模态生成与理解bilibiliA 类可核验real_case auto_approved 进入模型卡精选

GuanxingLu 使用 Kimi K2.5 处理研究分析和报告生成

GuanxingLu · Kimi K2.5

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.5 来源平台:github 最后复核:2026-06-26T23:14:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

GuanxingLu 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T23:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Researcher GuanxingLu built OpenPARL, a from-scratch reproduction of Kimi K2.5's PARL (Parallel Agent Reasoning Layer) Agent Swarm as described in the paper arXiv:2602.02276 'Kimi K2.5: Visual Agentic Intelligence'. The…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete open-source reproduction of the K2.5 PARL agent swarm architecture, including RL training pipelines, evaluation scripts, and ablation studies. The project demonstrates that K2.5's PARL agent swarm methodology…

模型作用:Kimi K2.5 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.5 provides the foundational research paper and architecture specification (arXiv:2602.02276) that OpenPARL reproduces. The PARL Agent Swarm design pattern from K2.5 is the core innovation being studied, includin…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a research reproduction, not direct K2.5 API usage. The actual models used are Qwen3-4B variants. The value is in validating and open-sourcing K2.5's architectural innovations for the broader research community.

原始记录:GuanxingLu reproduces Kimi K2.5 PARL Agent Swarm architecture on WideSearch with RL-trained orchestrator

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LQF-dev 使用 DeepSeek V4 Pro (Max) 处理智能体流程编排

LQF-dev · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LQF-dev 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LQF-dev built Zero Code Studio, an engineering-focused AI code generation platform based on Spring Boot + LangChain4j + LangGraph4j, using DeepSeek V4 Pro as the core LLM. The platform supports tool calling, Workflow or…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Full-stack AI code generation platform with dual-engine architecture (LangChain4j tool-calling mode and LangGraph4j workflow mode), covering generate → modify → preview → deploy. Includes DAG-based task planning, three-…

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro serves as the primary reasoning model powering the platform's code generation pipeline. It handles complex Vue engineering generation in reasoning mode and supports structured task planning via tool call…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Relatively small project (60 stars), single developer. Platform is functional but not widely adopted.

原始记录:Zero Code Studio: AI Code Generation Platform Powered by DeepSeek V4 Pro

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

madeye / maxlv 使用 DeepSeek V4 Pro (Max) 处理代码审查和测试生成

madeye / maxlv · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

madeye / maxlv 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:madeye built a Chinese A-share stock trading strategy dashboard called 'Silicon Civilization Stock Trade' (硅基文明消费股交易系统), focused on AI infrastructure supply chain stocks. The system uses DeepSeek V4 Pro (deepseek-v4-pro) for real-time trading signal generation and market analysis, with deepseek-v4-flash as the backtesting model. The system ingests market data via Tushare Pro + AkShare, generates DeepSeek-powered strategy signals, and supports TypeScript-based backtesting. The live dashboard is deployed at scs.maxlv.net.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production live dashboard (scs.maxlv.net) with real-time stock signals, target prices, and DeepSeek-powered strategy analysis for Chinese A-share AI infrastructure stocks. Includes backtesting engine, static snapshot ge…

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro is used as the primary strategy signal generation model, analyzing market data and producing trading recommendations. deepseek-v4-flash handles the computationally heavier backtesting tasks. The system c…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (46 stars), personal trading system. Financial claims should not be taken as investment advice.

原始记录:Silicon Civilization Stock Trade: AI-Powered A-Share Trading Signal System

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

rohitg00 使用 DeepSeek V4 Pro (Max) 处理智能体流程编排

rohitg00 · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

rohitg00 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:rohitg00 built AgentMemory, the top-ranked persistent memory system for AI coding agents (24K+ stars). The system uses DeepSeek V4 Pro as the recommended model for background memory compression and summarization tasks. …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Persistent memory server for AI coding agents that compresses old context via LLM summarization, supports multi-agent scoping (AGENT_ID), and provides recall/write APIs. Deployed as a sidecar service alongside coding ag…

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro is explicitly recommended as the optimal model for memory compression workloads, scoring within rounding error of Claude Sonnet 4.6 on summarization quality while costing $0.87/M output tokens vs $15/M (…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Model recommendation is based on developer benchmarks, not independent evaluation. Cost comparison may change with pricing updates.

原始记录:AgentMemory: Persistent Memory Compression for AI Coding Agents via DeepSeek V4 Pro

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yanqiyao1 使用 DeepSeek V4 Pro (Max) 处理软件工程任务执行

yanqiyao1 · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yanqiyao1 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:yanqiyao1 built SeekCode, a terminal-native code agent designed specifically for DeepSeek models. It defaults to deepseek-v4-pro as the primary model and deepseek-v4-flash as the fast companion. The tool is engineered f…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Full-featured terminal code agent (npm package: seekcode) with interactive mode, one-shot tasks, multi-provider support (deepseek, nvidia-nim, openrouter, novita, fireworks, sglang), local HTTP/SSE server mode, and Deep…

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro is the default and primary model for all code reasoning tasks. The entire context management, reasoning effort tuning, and thinking stream handling is purpose-built around DeepSeek V4 Pro's capabilities.…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (8 stars), early stage. Specific to DeepSeek model ecosystem.

原始记录:SeekCode: Terminal-Native Code Agent Built Around DeepSeek V4 Pro

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

maifeipin 使用 DeepSeek V4 Pro (Max) 处理知识检索和问答

maifeipin · DeepSeek V4 Pro (Max)

A
厂商:DeepSeek 模型:DeepSeek V4 Pro (Max) 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

maifeipin 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:maifeipin built Lite Agent, a lightweight AI assistant engine that connects to Feishu (WebSocket), Telegram, DingTalk, and WeChat Work. The system uses DeepSeek V4 Pro as its primary LLM backend and features a dynamic s…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production multi-channel AI assistant with 10+ built-in skills (RSS, billing, ops, security, blog, media, cloud backup), automatic OCR for image messages, cron-based scheduled tasks, multi-agent task orchestration, and …

模型作用:DeepSeek V4 Pro (Max) 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V4 Pro serves as the primary LLM backbone powering all natural language understanding, task decomposition, skill execution, and multi-turn conversation. The system leverages DeepSeek's thinking mode for complex…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Personal project (14 stars). Production deployment is for personal use. OCR endpoint points to private server.

原始记录:Lite Agent: Multi-Channel AI Assistant Engine with Feishu/Telegram/DingTalk Integration

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ai-co-id / ChatHermes 使用 Kimi K2 Thinking 处理软件工程任务执行

ai-co-id / ChatHermes · Kimi K2 Thinking

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 Thinking 来源平台:github 最后复核:2026-06-27T06:01:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ai-co-id / ChatHermes 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ChatHermes is presented as an open-source autonomous-agent SaaS that provisions dedicated or shared Hermes Agent instances, supports multi-model chat, and exposes real chat-time tools for research, drafting, coding, mon…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository states that it is the source for the publicly reachable managed service at chathermes.com and describes a production-style multi-tenant SaaS with 14 chat-time tools, hosted user-facing deployment, self-ho…

模型作用:Kimi K2 Thinking 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The project README explicitly says ChatHermes is built on Nous Research's Hermes Agent and Moonshot AI's Kimi K2 Thinking, and lists Kimi K2 / K2 Thinking among the supported multi-model chat options. Kimi K2 Thinking i…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-reported in the public GitHub README; both the repository and product URL were reachable during collection. The README names Kimi K2 Thinking explicitly, but does not expose private runtime logs proving…

原始记录:ai-co-id built the ChatHermes autonomous-agent SaaS on Hermes Agent and Kimi K2 Thinking

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yashb98 使用 Kimi K2.6 处理知识检索和问答

yashb98 · Kimi K2.6

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.6 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yashb98 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an AI talent intelligence agent that takes a job spec, discovers LinkedIn candidates through the user's real Chrome session (via Kimi WebBridge), scrapes profiles, scores them on a 5-gate weighted rubric (Skills 3…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete Bun+TypeScript application with web dashboard, SQLite persistence, Notion API integration, and Telegram notifications. The agent autonomously runs 3-5 SearchWeb queries to collect 20-30 LinkedIn profiles, sco…

模型作用:Kimi K2.6 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.6 serves as the core LLM powering the agent's autonomous planning, tool selection, candidate scoring logic, and personalized message drafting. It is accessed via @moonshot-ai/kimi-agent-sdk driving the local kim…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has low star count (new project). Uses Kimi WebBridge which requires user's real Chrome session running locally. No public deployment URL confirmed. Model usage is via Kimi CLI OAuth, not direct API.

原始记录:Talent Intelligence Agent — LinkedIn candidate discovery and scoring with Kimi K2.6

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

vinilpolepalli 使用 Kimi K2.6 处理软件工程任务执行

vinilpolepalli · Kimi K2.6

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.6 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

vinilpolepalli 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an autonomous AI-powered swing trading system that runs nightly on GitHub Actions: analyze 1,400+ stocks using Kimi K2.6 via NVIDIA NIM API, score each stock 0-100 by buy conviction based on price momentum, sector…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully autonomous trading pipeline: 4 PM ET daily, Kimi K2.6 analyzes 1,437 stocks (~6 hours). 8:30 AM ET, a secondary system reads the report and places proportional buys on Robinhood. 3:30 PM ET checks +7%/-5% exit t…

模型作用:Kimi K2.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.6 is the core analytical engine that evaluates each stock across 5 dimensions (price momentum, sector context, market conditions, technical setup, catalysts) and returns structured scores and signals. Accessed v…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Low star count. Automated trading involves financial risk. No confirmation of live trading performance or profitability. Uses NVIDIA NIM as inference provider, not direct Moonshot API.

原始记录:Phantom Trader — Autonomous AI swing trading bot analyzing 1400+ stocks with Kimi K2.6

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Yogurto710 使用 Kimi K2.6 处理软件工程任务执行

Yogurto710 · Kimi K2.6

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.6 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Yogurto710 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a command-line research assistant for public equities that produces sourced, dated briefs on specific stock questions (analyst research TICKER QUESTION) and full 8-section initiation reports with peer comps, valua…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Python CLI tool that produces equity research briefs (~$0.20-0.40/run, 2-3 min) and full initiation reports (~$0.60-0.80/run, 8-12 min). Output includes YAML frontmatter with ticker, date, language, model, and open qu…

模型作用:Kimi K2.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.6 is the default LLM for all research and initiation report generation. It natively writes in both English and Chinese on the first pass, preserving financial notation ($, tickers, URLs) byte-for-byte in Chinese…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Low star count. WeChat mini-app companion mentioned but not confirmed deployed. Uses Tavily for web search augmentation. Research output is not financial advice.

原始记录:Le Analyst — CLI equity research assistant with WeChat mini-app using Kimi K2.6

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Fuhad-Lab 使用 Kimi K2.6 处理代码审查和测试生成

Fuhad-Lab · Kimi K2.6

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.6 来源平台:github 最后复核:2026-06-26T15:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Fuhad-Lab 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T15:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Render-hosted backend service for the Blue Horizon e-learning platform that uses Kimi K2.6 to generate and edit interactive HTML learning modules from natural language prompts, and Playwright to visually test an…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A deployed Render web service with 5 endpoints: /health (status check), /generate (create module from prompt), /edit (modify existing module), /test (visual testing with Playwright screenshots), /start and /action (brow…

模型作用:Kimi K2.6 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.6 serves as the code generation engine, converting natural language teacher instructions into interactive HTML/CSS/JavaScript learning modules. Accessed via NVIDIA API. The model's coding capabilities enable it …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Low star count. Backend-only service with no public-facing demo URL. Uses NVIDIA API as inference provider. Blue Horizon platform deployment status unconfirmed.

原始记录:AI Module Tester — E-learning module generation and visual testing with Kimi K2.6

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Dolores Research 使用 MiniMax-M2.5 处理代码审查和测试生成

Dolores Research · MiniMax-M2.5

A
厂商:MiniMax 模型:MiniMax-M2.5 来源平台:github 最后复核:2026-06-26T23:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Dolores Research 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Open-source grounded reasoning agent that retrieves, extracts, computes, and validates answers from large enterprise document archives (697 Treasury Bulletin TXT files). Pipeline: question → grep retrieval → table parsi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:71.5% accuracy (176/246) on OfficeQA benchmark — #1 ranking on Sentient Arena Cohort 0 leaderboard with peak score 192.046. Achieved +3.7% accuracy over Claude Opus 4.5 at 1/500th the cost ($1.85 total / $0.0075 per que…

模型作用:MiniMax-M2.5 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.5 (via OpenRouter, Goose runtime) serves as the sole reasoning backbone — decomposes complex document questions, orchestrates tool calls (grep, Python computation), and produces validated answers. The 10B act…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 10 stars, by Dolores Research (doloresresearch.com). Benchmark results are on OfficeQA (Treasury Bulletin corpus), not general-purpose tasks. Uses Goose agent harness from Block Inc.

原始记录:Teller — #1 Sentient Arena grounded reasoning agent for enterprise document QA

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Rahul4u-stack (fintech PM) 使用 MiniMax-M2.5 处理研究分析和报告生成

Rahul4u-stack (fintech PM) · MiniMax-M2.5

A
厂商:MiniMax 模型:MiniMax-M2.5 来源平台:github 最后复核:2026-06-26T23:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Rahul4u-stack (fintech PM) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Personal AI agent that runs daily at 9:00 AM IST, uses minimax/minimax-m2.5:online on OpenRouter for LLM-driven web search of the last 24 hours of AI news, then curates 5 stories (research / products / industry / new mo…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Running production service since 2026-05-12. Delivers a single curated 5-story AI news brief to Telegram DM every morning. Uses Hermes 3 agent framework for scheduling, retry, and model routing.

模型作用:MiniMax-M2.5 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.5 (via OpenRouter with :online suffix) provides LLM-driven web search capability — the model autonomously searches the web for the latest AI news and selects/summarizes the 5 most relevant stories across 5 to…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Single user personal project. Model choice rationale: inexpensive on OpenRouter, default model in Hermes framework, quality sufficient for 5-story summarization. Author is a PM in fintech building an AI portfolio.

原始记录:Khabar — Daily AI news agent delivering curated briefs to Telegram via MiniMax M2.5

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Cuantox 使用 MiniMax-M2.5 处理软件工程任务执行

Cuantox · MiniMax-M2.5

A
厂商:MiniMax 模型:MiniMax-M2.5 来源平台:github 最后复核:2026-06-26T23:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Cuantox 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:CLI-based agentic AI coding companion that autonomously reads/writes files, executes shell commands, and fixes code in an auto-retry loop. Uses MiniMax M2.5 (minimaxai/minimax-m2.5) via NVIDIA API with OpenAI-compatible…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Fully functional CLI tool with agent loop: user prompt → model reasoning → tool calls (list files, read code, write files, execute commands) → error recovery → final output. Published on GitHub with MIT license, install…

模型作用:MiniMax-M2.5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.5 (via NVIDIA integrate API) is the sole reasoning and tool-calling engine — orchestrates multi-step coding tasks, generates code, parses tool call responses, and drives the auto-fix loop when commands fail. …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (1 star). Uses NVIDIA hosted endpoint (integrate.api.nvidia.com). Windows-optimized but cross-platform. Model access requires NVIDIA API key.

原始记录:Cuantox-AI — Agentic CLI coding companion powered by MiniMax M2.5 via NVIDIA API

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

re-cinq 使用 MiniMax-M2.5 处理医疗和生命科学分析

re-cinq · MiniMax-M2.5

A
厂商:MiniMax 模型:MiniMax-M2.5 来源平台:github 最后复核:2026-06-26T23:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

re-cinq 公开的医疗与生命科学案例,来源为 公开代码库,复核于 2026-06-26T23:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕医疗和生命科学分析的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Dockerized self-hosted inference server running MiniMax M2.5 (230B params, 10B active MoE) locally on NVIDIA DGX Spark (GB10 Grace Blackwell, 128GB unified memory). Uses llama.cpp with Q3_K_XL quantization (~101GB), ful…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Working Docker Compose deployment: single-command start (docker compose up -d), health check endpoint, OpenAI-compatible API. Tested with curl: model responds to chat completions requests with full reasoning capability.

模型作用:MiniMax-M2.5 在该案例中承担医疗和生命科学分析相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.5 provides the full reasoning backbone for the locally-hosted inference server. The MoE architecture (230B total, 10B active) enables efficient local deployment on consumer/prosumer hardware while maintaining…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:12 stars. Requires NVIDIA DGX Spark hardware (GB10 Grace Blackwell). Uses Unsloth GGUF quantization (UD-Q3_K_XL). Model download ~101GB. Not a 'use case' in the application sense but a verified deployment/configuration …

原始记录:Self-hosted MiniMax M2.5 inference server on NVIDIA DGX Spark with OpenAI-compatible API

已有真实案例 医疗与生命科学公开代码库A 类可核验real_case auto_approved

fciaf420 使用 MiniMax M2.7 处理研究分析和报告生成

fciaf420 · MiniMax M2.7

A
厂商:MiniMax 模型:MiniMax M2.7 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

fciaf420 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MoonBags is a Solana meme-token auto-trading bot that consumes real-time discovery streams from OKX smart-money signals and GMGN curated trenches, executes swaps via Jupiter Ultra, and uses MiniMax M2.7 as an optional L…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production Solana meme-token trading bot with Telegram control interface and web dashboard, where MiniMax M2.7 provides real-time, data-driven sell decisions based on multi-signal on-chain analysis. The bot operates i…

模型作用:MiniMax M2.7 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.7 serves as the cognitive exit advisor, analyzing complex multi-dimensional on-chain data (smart money flow, developer token holdings, holder profit/loss, price klines) every 30 seconds to make time-sensitive…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:M2.7 is an optional component; the bot also works with configurable trail/stop exits without LLM. Financial trading use case — disclaimer in repo states not financial advice. 34 GitHub stars as of June 2026.

原始记录:MoonBags: MiniMax M2.7 as LLM Exit Advisor for Solana Meme-Token Auto-Trading

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

bcefghj 使用 MiniMax M2.7 处理软件工程任务执行

bcefghj · MiniMax M2.7

A
厂商:MiniMax 模型:MiniMax M2.7 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

bcefghj 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Agent-Pilot is a multi-agent collaboration system built for the Feishu AI Campus Challenge. It uses 6 specialized AI Agents (Intent, Planner, Research, Writer, Review, Builder) with MiniMax M2.7 for tool calling, enabli…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully operational multi-agent pipeline with live demo (http://8.136.98.175), 50-page technical whitepaper, MCP Server for Cursor/Claude Desktop integration, and Flutter+WebSocket real-time dashboard. 14/14 PRD tests p…

模型作用:MiniMax M2.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.7 provides the tool-calling backbone for the multi-agent system — each of the 6 agents uses M2.7 for reasoning and tool invocation. The model's coding and agent capabilities enable structured task decompositi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Built for a specific competition (Feishu AI Campus Challenge). Live demo server at 8.136.98.175 may not be permanently available. 9 GitHub stars as of June 2026.

原始记录:Agent-Pilot: 6-Agent Multi-Agent Pipeline for Feishu IM-to-Presentation Automation Using MiniMax M2.7

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xiaoyza 使用 MiniMax M2.7 处理文档理解和结构化处理

xiaoyza · MiniMax M2.7

A
厂商:MiniMax 模型:MiniMax M2.7 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xiaoyza 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ModelDocUpgrade is a bidding document compliance review tool that uses MiniMax-M2.7-highspeed to automatically extract rules from a rule document (e.g., bidding requirements) and check them against a response document (…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A web-based document review tool (Node.js server + browser UI) that produces per-rule compliance assessments with four evidence states (met/partial/missing/risk) and an overall verdict (PASS/RISK/INCOMPLETE). Extracts r…

模型作用:MiniMax M2.7 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.7-highspeed is the core reasoning engine that reads structured document content, understands complex compliance rules from bidding documents, and evaluates whether response documents satisfy each rule. The mo…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Uses MiniMax-M2.7-highspeed variant specifically. Built for a niche but practical document compliance use case. Small repo, no stars recorded.

原始记录:ModelDocUpgrade: Word Document Rule Compliance Checker Using MiniMax-M2.7-highspeed

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DanoAndHolidays 使用 MiniMax M2.7 处理软件工程任务执行

DanoAndHolidays · MiniMax M2.7

A
厂商:MiniMax 模型:MiniMax M2.7 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DanoAndHolidays 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A multi-agent personalized learning platform built for the 15th China Software Cup (A3 track). The system uses multiple AI agents (profile construction, resource generation, path planning, intelligent tutoring, learning…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete AI-driven personalized learning platform with 5 collaborative agents covering the full learning lifecycle: student profiling, resource generation, learning path planning, real-time tutoring, and learning asse…

模型作用:MiniMax M2.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.7 serves as the foundational LLM powering all 5 learning agents. The model's coding and reasoning capabilities enable: (1) student profile construction from interaction data, (2) educational resource generati…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Competition submission for China Software Cup — may not be a production-deployed system. Uses Anthropic-compatible API format to call MiniMax. 2 GitHub stars.

原始记录:cnsoftbei: Multi-Agent AI Personalized Learning Platform for China Software Cup Using MiniMax-M2.7

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

EricSun0218 使用 MiniMax M2.7 处理软件工程任务执行

EricSun0218 · MiniMax M2.7

A
厂商:MiniMax 模型:MiniMax M2.7 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

EricSun0218 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A QQ group AI assistant bot that uses MiniMax M2 (M2.7 family) as its brain, combined with local code repository search (via ripgrep) and SQLite/FTS5 long-term memory. When mentioned in a QQ group, the bot answers code …

公开产物:A production QQ group bot with: (1) agentic code search using ripgrep with sandbox-locked repo access, (2) automatic memory extraction from conversations with deduplication, (3) manual memory management (remember/forget/list commands), (4) FTS5 Chinese bigram fuzzy recall for memory search, (5) multi-key rotation for rate limit resilience. Supports commands like '@bot 鉴权在哪实现的?' for code search and '@bot 记住:部署统一用 fly deploy' for memory storage.

模型作用:MiniMax M2.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2 (M2.7 family) acts as the reasoning brain with tool-calling capabilities: the model iteratively invokes rg_search, read_file, and list_dir tools to explore local codebases, synthesizes findings with long-term…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Uses NapCatQQ unofficial protocol bridge which violates Tencent ToS (repo warns to use alt account only). README references 'MiniMax M2' family rather than explicitly M2.7, but M2.7 is the latest model in the M2 family …

原始记录:qq-codememory-bot: QQ Group AI Bot with Local Code Search and Long-Term Memory Powered by MiniMax M2

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenAI (Inference Infrastructure Te… 使用 GPT-5 (high) 处理软件工程任务执行

OpenAI (Inference Infrastructure Team) · GPT-5 (high)

A
厂商:OpenAI 模型:GPT-5 (high) 来源平台:qq_news 最后复核:2026-06-26T15:14:08Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

OpenAI (Inference Infrastructure Team) 公开的代码代理与软件工程案例,来源为 qq_news,复核于 2026-06-26T15:14:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenAI's inference infrastructure team faced a load-balancing challenge with static token chunking. Codex (powered by GPT-5) analyzed weeks of production traffic data, wrote custom heuristics, and improved token generat…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Custom heuristic algorithm for dynamic load balancing was generated by Codex/GPT-5, resulting in >20% improvement in token generation throughput by replacing the previous static block-splitting approach.

模型作用:GPT-5 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5's reasoning enabled analysis of complex production traffic patterns, identification of inefficiencies in the static load-balancing approach, and generation of optimized heuristics — a self-improvement loop where A…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Evidence from Tencent News article (2026-04-28). This is an AI-self-improvement case: GPT-5/Codex optimized the inference infrastructure it runs on. The article explicitly states 'AI improved the system running itself' (字面意义上的'AI改进了跑自己的系统').

原始记录:OpenAI Codex (GPT-5) optimizes production inference infrastructure, boosting token generation speed 20%+

已有真实案例 代码代理与软件工程qq_newsA 类可核验real_case auto_approved 进入模型卡精选

Ridwan Oladipo, MD / MedNex AI 使用 GPT-5 (high) 处理研究分析和报告生成

Ridwan Oladipo, MD / MedNex AI · GPT-5 (high)

A
厂商:OpenAI 模型:GPT-5 (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Ridwan Oladipo, MD / MedNex AI 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a hospital-grade clinical decision support system for polypharmacy safety that unifies 170K+ DrugBank drug interactions with 77K RxNorm brand-to-ingredient mappings, performing real-time medication safety analysis…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production system deployed on AWS ECS Fargate at drug.mednexai.com with <200ms Tier-1 direct KB lookups, FAISS semantic retrieval over 170K vectors, and GPT-5-powered clinical reasoning generating color-coded severity f…

模型作用:GPT-5 (high) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5 (high) provides the clinical reasoning engine that classifies drug interactions into severity tiers, synthesizes polypharmacy risks across multiple drug pairs, and generates actionable clinical text output for pre…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Production deployment (drug.mednexai.com) was returning 503 at time of verification, README states it is 'available by request'. HuggingFace demo URLs referenced in README could not be independently verified due to netw…

原始记录:MedNex AI Drug Interaction Clinical Decision Support System

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

agruai 使用 GPT-5 (high) 处理智能体流程编排

agruai · GPT-5 (high)

A
厂商:OpenAI 模型:GPT-5 (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

agruai 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a production-ready multi-agent book-writing system using CrewAI that transforms author briefs into complete manuscripts with 5 specialized AI agents (Concept Generator, Outliner, Writer, Editor, Continuity Checker…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Full-stack AI book-writing platform with FastAPI backend, Next.js 14 frontend, PostgreSQL 16, Redis clustering, Kubernetes deployment on GKE, and CrewAI multi-agent orchestration. Supports sub-300ms SSE streaming, 40% A…

模型作用:GPT-5 (high) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5 is the primary LLM powering the 5 specialized CrewAI agents (via LiteLLM routing with fallback chains to Claude Sonnet 4 and Gemini 2.5 Pro). GPT-5 handles the core content generation, editing, and continuity chec…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 37 stars. Production deployment status not confirmed. GPT-5 is one of multiple supported models with fallback chains.

原始记录:Book.ai Multi-Agent AI Book Writing System

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

IuriiD 使用 GPT-5 (high) 处理研究分析和报告生成

IuriiD · GPT-5 (high)

A
厂商:OpenAI 模型:GPT-5 (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

IuriiD 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an AI-powered UGC advertisement video cloning system that analyzes successful UGC ad videos using Gemini for visual analysis, then generates new promotional videos using GPT-5 for text/prompt generation and Sora 2…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Full-stack NestJS + React application that analyzes reference UGC ad videos (visual style, messaging, pacing, engagement techniques), generates AI-powered text-to-video prompts combining original insights with product d…

模型作用:GPT-5 (high) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5 handles the text generation and prompt engineering pipeline — synthesizing visual analysis insights with product information to create optimized video generation prompts, and performing content moderation on gener…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Uses GPT-5 via Laozhang API proxy, not direct OpenAI API. Repo has 34 stars. POC status.

原始记录:viral2viral AI UGC Advertisement Video Generator

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

tahangz 使用 GPT-5 (high) 处理多模态内容处理

tahangz · GPT-5 (high)

A
厂商:OpenAI 模型:GPT-5 (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tahangz 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a full-stack voice assistant application integrated with Vapi and GPT-5 that collects user information via form, conducts real-time voice-based conversations with configurable objectives (e.g., evaluating job/prog…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Working full-stack application with React frontend and Flask backend. Users fill in personal details, engage in voice-based conversation with GPT-5-powered Vapi assistant configured with specific screening objectives, a…

模型作用:GPT-5 (high) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5 powers the conversational intelligence of the Vapi voice assistant, enabling real-time natural language understanding, context-aware responses, configurable objective-based dialogue, and post-call summary generati…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 10 stars. No evidence of production deployment beyond the open-source project. Vapi integration adds a dependency layer.

原始记录:Vapi-Powered Voice Assistant with GPT-5 for Job Screening

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

matsumotokohei 使用 Kimi K2.6 处理可玩交互原型构建

matsumotokohei · Kimi K2.6

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.6 来源平台:github 最后复核:2026-06-26T16:06:42.969341+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

matsumotokohei 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-26T16:06:42.969341+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Japanese developer (matsumotokohei) used Kimi K2.6 as the primary coding model via OpenCode to build a Street Fighter 6-style 2D fighting game in Godot 3.6 GDScript. The project features 7 specialized AI sub-agents (m…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete playable Godot 3.6 fighting game project with 25+ GDScript files totaling ~350KB of game logic, including BattleManager (20KB), CPUController (36KB), CharacterData (35KB), Fighter (49KB), FighterVisual (56KB)…

模型作用:Kimi K2.6 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.6 served as the sole coding model powering all 7 AI agents in the OpenCode workflow: the manager agent handled requirements and task delegation, the developer agent implemented GDScript code, the reviewer agent …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 0 stars as of collection date; model is configured at OpenCode system level not in repo config files, but repo description explicitly states 'Model Kimi K2.6'; project appears to be a personal hobby project by …

原始记录:Kimi K2.6 builds a SF6-style fighting game with 20 characters via OpenCode multi-agent workflow

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ZacharyZhang-NY 使用 Kimi K2.6 处理软件工程任务执行

ZacharyZhang-NY · Kimi K2.6

A
厂商:Kimi / Moonshot AI 模型:Kimi K2.6 来源平台:github 最后复核:2026-06-26T16:06:42.969375+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ZacharyZhang-NY 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:06:42.969375+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A developer (ZacharyZhang-NY) ran the same detailed PRD (FlowBoard - a team project management board with drag-and-drop Kanban, Sprint planning, and Issue tracking) through 4 different AI coding models: GLM 5.1, Kimi K2…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete Next.js 16 application codebase in the Kimi-K2.6 directory including app routes, backend API, database schema, components, hooks, types, configuration files (next.config.ts, drizzle.config.ts, tsconfig.json),…

模型作用:Kimi K2.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.6 generated the entire full-stack application from a single PRD prompt, including frontend components with IBM Carbon Design System, backend API routes, database schema with Drizzle ORM, authentication with Bett…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 2 stars; this is a model comparison project so Kimi K2.6 output is presented alongside other models; the PRD explicitly targets Kimi Code as the coding agent; Docker deployment suggests the generated code is fu…

原始记录:Kimi K2.6 builds a full-stack FlowBoard project management app in the same PRD challenge as Claude Opus 4.6, GPT-5.4, and GLM 5.1

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

loke-jad 使用 MiniMax-M2.5 处理软件工程任务执行

loke-jad · MiniMax-M2.5

A
厂商:MiniMax 模型:MiniMax-M2.5 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

loke-jad 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Claude Agent SDK wired to a local llama-server running MiniMax-M2.5 (quantized, ~101GB) via LiteLLM on an NVIDIA DGX Spark (128GB unified memory). Includes MCP (Model Context Protocol) gateway for tool integration and s…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Working local coding agent with full Claude Code compatibility. Detailed optimization plan documented (OPTIMIZATION.md): expected 10-20x prompt processing speedup via persistent KV cache warming. Agent runs fully locall…

模型作用:MiniMax-M2.5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M2.5 (via Unsloth GGUF quantization, llama-server, and LiteLLM proxy) serves as the sole LLM backbone for the coding agent. The model handles code generation, tool orchestration, multi-turn conversation, and rea…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Project marked as 'Dormant' by author. No README.md file — evidence comes from OPTIMIZATION.md and GitHub page metadata. 0 stars. Strong technical evidence of real deployment (detailed VRAM measurements, cache profiling…

原始记录:Claude Agent SDK + local MiniMax M2.5 via LiteLLM — MCP gateway coding agent with safety hooks

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

haydenkz 使用 GPT-5 Codex (high) 处理可玩交互原型构建

haydenkz · GPT-5 Codex (high)

A
厂商:OpenAI 模型:GPT-5 Codex (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

haydenkz 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发者使用 GPT-5 Codex 自动生成一个基于 Next.js 15 + React 19 + TypeScript 的 Cartpole(倒立摆)物理模拟游戏,从零开始全代码由模型生成,无需手写代码。

公开产物:完整可运行的 Cartpole 物理模拟游戏,包含前端渲染、物理引擎逻辑、交互控制,使用 Next.js 框架打包为可部署应用。README 明确标注 'GPT-5 Codex made this'。

模型作用:GPT-5 Codex 全权生成了项目全部代码,包括物理模拟逻辑、React 组件、TypeScript 类型定义和构建配置,开发者未手动编写任何代码。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 表明 'don't expect bug fixes or upgrades',项目为一次性生成产物,后续维护由用户自行处理。

原始记录:GPT-5 Codex 生成 TypeScript Cartpole 物理模拟游戏

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

pkostelnik 使用 GPT-5 Codex (high) 处理软件工程任务执行

pkostelnik · GPT-5 Codex (high)

A
厂商:OpenAI 模型:GPT-5 Codex (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

pkostelnik 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发者使用 GPT-5 Codex 构建一个面向装备潜水员的完整潜水日志管理平台 DiveLog Studio,涵盖登录注册、潜水记录管理、设备管理、潜点数据库、社区分享等功能。

公开产物:功能完备的全栈 Web 应用(Next.js 16.2 + TypeScript 6.0 + React 19.2 + Tailwind CSS 4.2),支持双语(德/英)、暗色模式、Microsoft Teams 集成、Gravatar 头像、5 种海洋主题配色,已部署至 divelog.copilot.ovh 提供在线演示。

模型作用:GPT-5 Codex 生成了完整的前后端代码,包括组件架构、认证系统、国际化方案、响应式设计和 PWA 配置,GitHub 仓库以 gpt-5-codex 为 topic 标签。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目由单人开发者使用 Codex 构建,后续维护可持续性取决于开发者个人投入。

原始记录:GPT-5 Codex 构建 DiveLog Studio 全功能潜水日志 Web 平台

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenAI 使用 GPT-5 Codex (high) 处理浏览器 3D 世界构建

OpenAI · GPT-5 Codex (high)

A
厂商:OpenAI 模型:GPT-5 Codex (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenAI 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:OpenAI 团队使用 Codex 构建 ImageGenCam:一个基于 Raspberry Pi Zero 2 W 的可 DIY 数字相机项目,包含 3D 打印外壳设计、相机固件、图像生成管线和手机端配套 Web 应用。

公开产物:完整的硬件+软件项目:Raspberry Pi 相机固件(Python)、3D 打印外壳模型、基于 Codex 桌面应用的图像生成流程、手机端照片管理和 Prompt 更新 Web 应用。提供从零件清单到完整组装的全流程教程。

模型作用:Codex 被用于生成相机控制软件、图像生成管线代码、配套 Web 应用,以及协助 3D 模型定制。项目 README 明确标注 'a digital camera you can build yourself with Codex'。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:OpenAI 官方示范项目,展示了 Codex 在硬件/IoT 领域的跨模态生成能力。

原始记录:OpenAI 使用 Codex 构建 ImageGenCam 数字相机硬件项目

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenAI 使用 GPT-5 Codex (high) 处理软件工程任务执行

OpenAI · GPT-5 Codex (high)

A
厂商:OpenAI 模型:GPT-5 Codex (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenAI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:OpenAI 团队使用 Codex 构建 Euphony:一个浏览器端的 Codex 会话和 Harmony 对话可视化工具,支持 JSONL 文件加载、会话时间线渲染、翻译、元数据检查和可嵌入 Web 组件。

公开产物:完整的 Web 可视化应用,包含可嵌入 Web Components、JMESPath 过滤、网格/编辑器模式、多源数据加载(剪贴板/本地文件/HTTP URL),已部署至 openai.github.io/euphony/。400+ GitHub Stars。

模型作用:Codex 生成了 Euphony 的核心 Web Components、会话解析逻辑、Harmony 对话渲染器和部署配置,使 Codex 会话数据可被人类直观审查和分析。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:OpenAI 官方项目,服务于 Codex 生态系统的可观测性需求。

原始记录:OpenAI 使用 Codex 构建 Euphony Codex 会话可视化工具

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenAI 使用 GPT-5 Codex (high) 处理软件工程任务执行

OpenAI · GPT-5 Codex (high)

A
厂商:OpenAI 模型:GPT-5 Codex (high) 来源平台:github 最后复核:2026-06-26T16:00:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenAI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T16:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:OpenAI 团队使用 Codex CLI 构建一个自动化迁移工具包,帮助开发者将遗留的 Completions/Chat Completions API 应用迁移到统一的 Responses API,包含自动检测、代码编辑、测试运行和 PR 生成。

公开产物:完整的 Bash 迁移工具包:自动检测遗留 API 调用、提出并应用代码修改、更新导入和请求/响应结构、运行测试和 lint、创建干净的 git 分支和可选 PR。附带演示视频。105+ GitHub Stars。

模型作用:Codex CLI 被用于自动化代码迁移流程,从分析到编辑到测试再到 PR 创建,展示了 Codex 在大规模代码重构场景中的端到端自动化能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:迁移工具本身服务于 OpenAI API 生态,实际迁移效果取决于目标代码库的复杂度。

原始记录:OpenAI 使用 Codex 构建 Completions→Responses API 迁移工具包

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

liangdabiao (SeekMoney-ai) 使用 GLM-4.7 处理软件工程任务执行

liangdabiao (SeekMoney-ai) · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:github 最后复核:2026-06-27T00:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

liangdabiao (SeekMoney-ai) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T00:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SeekMoney-ai is a web application that helps entrepreneurs discover business opportunities from social media by automatically identifying core user pain points. The system collects data from 8 platforms (Douyin, TikTok,…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production web application that processes social media data across 8 platforms, performs AI-driven semantic clustering and deep pain point analysis, outputs structured market opportunity reports with priority scoring …

模型作用:GLM-4.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7's thinking model (thinking/reasoning mode) is the core analytical engine that performs deep multi-layer analysis of user pain points extracted from social media conversations. The model goes beyond surface-level…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Open-source GitHub repo with detailed Chinese README. Uses OpenAI-compatible API pointing to GLM endpoints. This is a third-party developer's product, not an official Zhipu customer story. The project explicitly states it uses 'openai兼容/GLM-4.7 思考模型' (GLM-4.7 thinking model).

原始记录:SeekMoney-ai: Multi-platform social media business opportunity finder powered by GLM-4.7 thinking model for deep user pain point analysis

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

JonniTech 使用 GLM-4.7-Flash 处理知识检索和问答

JonniTech · GLM-4.7-Flash

A
厂商:Z AI / GLM 模型:GLM-4.7-Flash 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

JonniTech 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建高保真 Perplexity AI 搜索界面克隆,核心 RAG 流水线使用 GLM-4.7-Flash 作为 LLM:用户查询→SerpAPI 获取实时网页搜索结果→上下文提取→GLM-4.7-Flash 生成流式回答并附带引用来源。

公开产物:完整的 RAG 搜索引擎应用,支持实时网页搜索、流式回答、Markdown 渲染、代码高亮、引用来源展示、用户认证(Clerk)和对话历史持久化。React 19 + TypeScript + Tailwind CSS 技术栈。

模型作用:GLM-4.7-Flash 是 RAG 流水线中的核心 LLM,负责接收 SerpAPI 搜索结果构建的上下文,生成准确、带引用的自然语言回答。整个应用的功能正确性直接依赖 GLM-4.7-Flash 的推理和生成能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目为个人 demo 项目,stars 较少;使用 GLM-4.7-Flash(免费快速版)而非 GLM-4.7 主版本。

原始记录:Perplexity Clone: 基于 GLM-4.7-Flash 的 RAG 搜索引擎

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

applex250 使用 GLM-4.7 处理文档理解和结构化处理

applex250 · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

applex250 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:使用智谱 GLM-4.7 构建英文学术文献 Markdown 翻译工具,支持 PDF/Word/PPT/图片等多格式输入(通过 MinerU 转换为 Markdown),智能分段、并发翻译、术语一致性保持、表格/公式/代码块过滤与保留,以及 Web 可视化界面。

公开产物:完整的翻译工具,支持命令行和 Web 界面(拖拽上传、任务队列、实时日志),输出高质量中文学术文献 Markdown。过滤特殊元素后翻译可节省 50%+ Token 消耗。v2.1 版本集成 MinerU API 支持多格式直接翻译。

模型作用:GLM-4.7 是翻译流水线的核心 LLM,负责将英文学术文本翻译为中文,同时保持术语一致性和学术语言准确性。智能分段后并发调用 GLM-4.7 API 进行翻译。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目为个人开源项目;.env 示例中默认使用 glm-4 而非 glm-4.7,但 README 和项目描述明确标注基于 GLM-4.7。

原始记录:mtrans: 基于 GLM-4.7 的英文学术文献 Markdown 翻译工具

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ChuanMeng (SIGIR 2026 研究团队) 使用 GLM-4.7-Flash 处理研究分析和报告生成

ChuanMeng (SIGIR 2026 研究团队) · GLM-4.7-Flash

A
厂商:Z AI / GLM 模型:GLM-4.7-Flash 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ChuanMeng (SIGIR 2026 研究团队) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:SIGIR 2026 论文《Revisiting Text Ranking in Deep Research》使用 GLM-4.7-Flash (30B) 作为深度研究 Agent,在 BrowseComp-Plus 数据集上评估 5 种检索器(BM25、SPLADE-v3、RepLLaMA、Qwen3-Embedding-8B、ColBERTv2)和 3 种重排器(monoT5-3B、RankLLaMA-7B、Rank1-7B)的文本排序性能。

公开产物:发表于 SIGIR 2026(第 49 届 ACM SIGIR 信息检索国际会议)的研究论文,提供了深度研究场景下文本排序方法的全面评估。发布了 BrowseComp-Plus 段落语料库、所有检索器索引和完整执行轨迹数据。

模型作用:GLM-4.7-Flash (30B) 作为两个深度研究 Agent 之一(与 gpt-oss-20b 并列),通过本地 vLLM 服务器部署运行,在多轮检索-重排序循环中执行信息检索和推理任务,是论文实验评估的核心 Agent 模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:使用 GLM-4.7-Flash (30B) 而非 GLM-4.7 主版本;属于学术研究实验而非生产应用,但有完整的论文发表和开源代码,证据等级高。

原始记录:SIGIR 2026: GLM-4.7-Flash 作为深度研究 Agent 进行文本排序评估

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

o1ie12 使用 Muse Spark 处理多模态内容处理

o1ie12 · Muse Spark

A
厂商:Meta / Llama 模型:Muse Spark 来源平台:GitHub 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

o1ie12 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Generate a complete single-file marketing website for 'Northstar Consulting' from natural language prompts — including 6 pages with hash routing, GSAP fade transitions, Lenis smooth scroll, Three.js blob backgrounds, cu…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A 45KB single index.html file containing a fully functional 6-page marketing studio website with smooth animations, mobile hamburger menu, work filters, expandable service cards, scrollable modals, and form validation. …

模型作用:Muse Spark 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Muse Spark generated the entire website from 5-6 prompts with zero manual code edits. The user provided a detailed prompt specifying design tokens (ivory/charcoal/sage/sand/terracotta palette, Fraunces + Inter fonts), i…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Single user project, not production-scale deployment. Demo site hosted on GitHub Pages free tier. Compared to Gemini in same test.

原始记录:o1ie12 built a complete marketing website using only Muse Spark prompts — no manual code edits

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

sci-freak 使用 Muse Spark 处理研究分析和报告生成

sci-freak · Muse Spark

A
厂商:Meta / Llama 模型:Muse Spark 来源平台:GitHub 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

sci-freak 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Design and build a lightweight cross-platform desktop/mobile application for PhD students to track experiments, papers, and milestones — with offline support, JSON storage, and cloud sync via Google Drive.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:PhD Tracker Lite: a Tauri 2 application (Rust backend + React frontend) with compiled release binaries for Windows (MSI) and Android (APK). Features include experiment tracking, paper management, milestone logging, offl…

模型作用:Muse Spark 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Muse Spark was used as the primary development assistant throughout the project lifecycle — from architecture decisions (Tauri 2 over Electron for native performance) to implementing the Rust backend, React frontend, an…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:First release not code-signed. Small-scale personal project. No formal metrics on how much of the code was AI-generated vs human-written.

原始记录:sci-freak built a cross-platform PhD progress tracker app (Windows + Android) using Muse Spark

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

compnew2006 使用 Muse Spark 处理研究分析和报告生成

compnew2006 · Muse Spark

A
厂商:Meta / Llama 模型:Muse Spark 来源平台:GitHub 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

compnew2006 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a self-hosted Go API that wraps Meta AI (Muse Spark) into an OpenAI-compatible endpoint, paired with a React 19 frontend ('SMART Studio') offering 11 specialized AI studios for branding, photography, video, voiceo…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Two integrated components: (1) metaai-go — a high-performance Go API wrapper for Meta AI using cookie authentication, exposing OpenAI-compatible /v1/chat/completions endpoint; (2) smart-studio — a React 19 frontend with…

模型作用:Muse Spark 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Muse Spark is the core LLM powering all AI features across SMART Studio's 11 studios and the multi-platform chat interface. The Go API wrapper specifically targets Meta AI's Muse Spark session-based architecture for pro…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Uses cookie-based authentication (undocumented Meta AI session mechanism). Educational purposes only per project disclaimer. May break if Meta changes UI/API.

原始记录:compnew2006 built a Go API + React 19 frontend to run Muse Spark across Telegram, WhatsApp, and Web

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

dyagz 使用 Muse Spark 处理智能体流程编排

dyagz · Muse Spark

A
厂商:Meta / Llama 模型:Muse Spark 来源平台:GitHub 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

dyagz 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a browser-backed proxy framework that exposes Meta AI's Muse Spark through both OpenAI-compatible /v1/chat/completions and Anthropic-compatible /v1/messages API frontends, with optional broker model integration fo…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Node.js proxy server with three-layer architecture: frontends (Anthropic + OpenAI format adapters), browser providers (meta.js for Meta AI session automation), and broker providers (zai.js for Z.ai GLM tool execution)…

模型作用:Muse Spark 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Muse Spark serves as the primary text generation backend accessed through the browser provider layer. The proxy automates browser interaction with Meta AI to provide structured API access to Muse Spark's capabilities fo…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Browser-backed integration is inherently fragile — UI selectors, auth flows, and anti-bot behavior can change without notice. Single-session only. Relies on undocumented Meta AI web interface.

原始记录:dyagz built a browser-backed LLM proxy exposing Muse Spark through OpenAI and Anthropic compatible APIs

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

red-teamIA 使用 Muse Spark 处理代码审查和测试生成

red-teamIA · Muse Spark

A
厂商:Meta / Llama 模型:Muse Spark 来源平台:GitHub 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

red-teamIA 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Conduct exploratory red team testing of Muse Spark via WhatsApp to probe safety boundaries — specifically testing whether the model hallucinates internal capabilities (e.g., claiming to have human review queues, moderat…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A structured bug report documenting that Muse Spark hallucinated possessing 'a critical human review queue' for fraud — an internal tool/capability the model does not actually have. The report includes: reproduction env…

模型作用:Muse Spark 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:The testing directly engaged Muse Spark's safety mechanisms and content moderation behavior, revealing that the model fabricated internal processes it attributed to Meta — a hallucination of capability that could mislea…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Portuguese-language report. Single-session test, not reproduced with other prompts per author. Bug report targets Meta's model specifically — could be used to highlight safety weaknesses.

原始记录:red-teamIA conducted systematic red team testing of Muse Spark safety on WhatsApp, documenting hallucination bugs

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Sourcegraph 使用 Claude 2.0 处理软件工程任务执行

Sourcegraph · Claude 2.0

A
厂商:Anthropic / Claude 模型:Claude 2.0 来源平台:official_web 最后复核:2026-06-27T10:11:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Sourcegraph 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T10:11:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Sourcegraph integrated Claude 2 into Cody, its AI coding assistant for helping developers write, fix, and maintain code. The concrete task was codebase-aware developer assistance: answering user queries with more reposi…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic's Claude 2 launch post states that Cody used Claude 2's improved reasoning to give more accurate answers to user queries while passing along more codebase context with up to a 100K-token context window. Source…

模型作用:Claude 2.0 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2 supplied the large-context LLM reasoning layer inside Cody, enabling Sourcegraph to include more repository context in prompts, answer code questions more accurately, and use Claude 2's more recent training dat…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Primary evidence is Anthropic's official Claude 2 launch post with a direct Sourcegraph customer quote; artifact URL is the current public Cody documentation/product page and may reflect newer model support rather than …

原始记录:Sourcegraph Cody used Claude 2 for codebase-aware AI coding assistance

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Sourcegraph 使用 Claude 2.1 处理知识检索和问答

Sourcegraph · Claude 2.1

A
厂商:Anthropic / Claude 模型:Claude 2.1 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Sourcegraph 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Sourcegraph added Claude 2.1 as a selectable model for their Cody AI coding assistant, enabling developers to use Claude 2.1 for code completion, explanation, and generation within the Sourcegraph platform. The PR merge…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Claude 2.1 became available as 'Claude 2.1 Preview' (model id: anthropic/claude-2.1) in Cody's model selector, alongside Claude 2.0 and GPT-4 Turbo. Users could leverage Claude 2.1's 200K context window for analyzing en…

模型作用:Claude 2.1 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2.1 provided a 200K token context window that allowed Cody to analyze large codebases in full, along with reduced hallucination rates for more accurate code suggestions and explanations.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Claude 2.1 has since been superseded by Claude 3.x and later models. Sourcegraph now supports newer Claude models.

原始记录:Sourcegraph Cody AI coding assistant integrates Claude 2.1 for code intelligence

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MLflow (Databricks) 使用 Claude 2.1 处理代码审查和测试生成

MLflow (Databricks) · Claude 2.1

A
厂商:Anthropic / Claude 模型:Claude 2.1 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MLflow (Databricks) 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MLflow implemented chat and chat streaming for Anthropic models including Claude 2.1 in its model deployment server, allowing ML engineers to deploy and serve Claude 2.1 as a production endpoint with OpenAI-compatible A…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:MLflow's deployment server gained full support for Claude 2.1 chat API with both non-streaming and streaming responses, accessible via REST endpoint (POST /endpoints/anthropic/invocations) with YAML config specifying 'c…

模型作用:Claude 2.1 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2.1's chat API enabled MLflow to provide enterprise-grade deployment infrastructure for organizations wanting to integrate Claude 2.1 into their ML pipelines with cost tracking, monitoring, and A/B testing capabi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:MLflow now supports newer Claude models. This PR specifically targeted Claude 2.1 as the primary model.

原始记录:MLflow adds Claude 2.1 chat and streaming support to its deployment server

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Graphlit 使用 Claude 2.1 处理研究分析和报告生成

Graphlit · Claude 2.1

A
厂商:Anthropic / Claude 模型:Claude 2.1 来源平台:web 最后复核:2026-06-26T23:55:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Graphlit 公开的研究与报告生成案例,来源为 web,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Graphlit, a content management platform, integrated Claude 2.1 to summarize multi-page ArXiv research papers into concise paragraphs, demonstrating its model-agnostic approach by comparing Claude 2.1's summarization qua…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Claude 2.1 produced significantly improved summaries of academic papers compared to Claude 2.0, with Graphlit's blog post showing Claude 2.1 generating accurate, concise summaries of the 'unification of LLMs and knowled…

模型作用:Claude 2.1 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2.1's improved comprehension and reduced hallucination rates made it the preferred model for Graphlit's content summarization pipeline, providing more accurate and reliable summaries of complex technical document…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Graphlit is a smaller platform. The blog post is from the day of Claude 2.1 launch (Nov 21, 2023) and serves as a product demonstration rather than a large-scale deployment.

原始记录:Graphlit uses Claude 2.1 for model-agnostic ArXiv paper summarization

已有真实案例 研究与报告生成webA 类可核验real_case auto_approved 进入模型卡精选

BerriAI (LiteLLM) 使用 Claude 2.1 处理软件工程任务执行

BerriAI (LiteLLM) · Claude 2.1

A
厂商:Anthropic / Claude 模型:Claude 2.1 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

BerriAI (LiteLLM) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:BerriAI's LiteLLM (51,000+ stars) added support for Claude 2.1, enabling organizations to access Claude 2.1 through a unified OpenAI-compatible proxy API with cost tracking, load balancing, and rate limiting. Real users…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:LiteLLM provided a production-ready proxy for Claude 2.1 with OpenAI-compatible API format, enabling organizations to swap between Claude 2.1 and other models without changing their application code. Multiple bug report…

模型作用:Claude 2.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2.1 was one of the key models supported by LiteLLM's multi-provider gateway, with specific model ID mapping and Anthropic API format handling that enabled thousands of organizations to adopt Claude 2.1 in product…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:LiteLLM is a proxy/gateway tool rather than a direct end-user application. The evidence comes from bug reports which confirm real usage but don't describe specific business outcomes.

原始记录:BerriAI LiteLLM enables multi-provider access to Claude 2.1 through unified proxy

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

atisharma (Chasm Engine) 使用 Claude 2.1 处理可玩交互原型构建

atisharma (Chasm Engine) · Claude 2.1

A
厂商:Anthropic / Claude 模型:Claude 2.1 来源平台:github 最后复核:2026-06-26T23:55:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

atisharma (Chasm Engine) 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-26T23:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Chasm Engine, a generative text adventure game engine (70 stars), specifically requested and integrated Claude 2.1 API access for narrative generation, scene creation, and character dialogue in an interactive fiction ga…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Chasm Engine used Claude 2.1 as a recommended model for generating game narratives, character interactions, and world descriptions. The engine leverages Claude 2.1's instruction following and creative writing capabiliti…

模型作用:Claude 2.1 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2.1's strong instruction following, creative writing ability, and reduced hallucination rates made it suitable for generating coherent game narratives that maintain consistency across multiple turns of player int…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Chasm Engine is a small indie project (70 stars). The developer specifically requested Claude 2.1 API access, indicating active preference for this model version.

原始记录:Chasm Engine uses Claude 2.1 for generative text adventure narrative generation

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ndom91 使用 GLM-4.7 处理软件工程任务执行

ndom91 · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:github 最后复核:2026-06-27T00:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ndom91 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T00:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer ndom91 deployed GLM-4.7-Flash locally on a Framework Mainboard with AMD Ryzen AI Max 395+ (Strix Halo) and 128GB RAM, running the model via llama.cpp with ROCm/Vulkan GPU acceleration in a Docker container. Th…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working local deployment of GLM-4.7-Flash on consumer AMD hardware (Framework Mainboard, Ryzen AI Max 395+, 128GB) achieving daily-use performance for coding assistance. The developer reports near-Claude-Code responsi…

模型作用:GLM-4.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7-Flash's efficiency and performance at local inference enables a fully offline, private coding assistant experience. The model's quality is high enough that the developer considers it a viable alternative to clou…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source GitHub repo with Docker Compose setup and clear documentation. Developer's subjective speed comparison to Claude Code. Hardware-specific (Strix Halo with ROCm support). This is an individual developer's depl…

原始记录:GLM-4.7-Flash on AMD Strix Halo: Local deployment achieving near-Claude-Code speed via llama.cpp ROCm with Vulkan acceleration

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

stefandevo 使用 GLM-4.7 处理软件工程任务执行

stefandevo · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:github 最后复核:2026-06-27T00:30:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

stefandevo 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T00:30:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer stefandevo built glm-acp-agent, a TypeScript implementation of the Agent Client Protocol (ACP) that uses Zhipu AI's GLM model family (specifically GLM-4.7 and GLM-5.1) as its reasoning core. The agent connects…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An open-source TypeScript agent (23 GitHub stars) that provides ACP-compatible integration of GLM-4.7 into developer IDEs, enabling real-time streaming responses, file system operations, terminal commands, and web acces…

模型作用:GLM-4.7 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-4.7 serves as the primary reasoning engine for the agent's decision-making and tool-calling capabilities. The model's instruction following, function calling support, and long-context handling enable the agent to au…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source GitHub repo with 23 stars. Written in TypeScript, specifically designed for Z.AI GLM Coding Plan. Developer explicitly chose GLM-4.7 for its agentic and tool-coding capabilities. Third-party developer projec…

原始记录:glm-acp-agent: TypeScript ACP agent using GLM-4.7 as reasoning core for IDE-integrated coding assistance with file system and terminal tools

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Ericklzam / Pivot Point Orthopaedic… 使用 Claude 3.5 Haiku 处理代码审查和测试生成

Ericklzam / Pivot Point Orthopaedics voice agent project · Claude 3.5 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3.5 Haiku 来源平台:github 最后复核:2026-06-27T00:22:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Ericklzam / Pivot Point Orthopaedics voice agent project 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T00:22:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Automated stress-testing of a healthcare voice agent (Pivot Point Orthopaedics) by using Claude 3.5 Haiku as an AI adversary that generates chaotic, unpredictable patient personas (confused patients, emergencies, distra…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working automated testing suite that calls the target voice agent via Retell AI WebSockets, with Claude 3.5 Haiku acting as the adversarial caller. Includes Loom video walkthrough and Google Drive video showing bug di…

模型作用:Claude 3.5 Haiku 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Haiku serves as the AI brain that generates and role-plays diverse, unpredictable patient personas in real-time via WebSockets, probing the target voice agent's guardrails and state management. Its low latenc…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small personal project (0 stars), single developer. Evidence is strong (code + Loom video) but limited in scale.

原始记录:Claude 3.5 Haiku as adversarial patient persona generator for voice agent stress testing

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

saitarrun 使用 Claude 3.5 Haiku 处理知识检索和问答

saitarrun · Claude 3.5 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3.5 Haiku 来源平台:github 最后复核:2026-06-27T00:22:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

saitarrun 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T00:22:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A GenAI developer productivity assistant that indexes local codebases with language-aware chunking, stores embeddings in Pinecone, and answers natural-language questions about the code using Claude 3.5 Haiku via Amazon …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Deployable CLI tool and serverless backend (AWS Lambda + API Gateway + Pinecone) that returns source-cited answers with file paths and relevance scores for code-related questions.

模型作用:Claude 3.5 Haiku 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Haiku processes RAG-retrieved code context and generates natural-language answers with source citations. Its low latency and cost efficiency make it suitable for a serverless, on-demand query architecture. Us…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small personal project (0 stars). Deployed on AWS Bedrock (us-east-1). Architecture is well-documented.

原始记录:Serverless RAG system for semantic code search using Claude 3.5 Haiku on AWS Bedrock

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

jordann6 使用 Claude 3.5 Haiku 处理研究分析和报告生成

jordann6 · Claude 3.5 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3.5 Haiku 来源平台:github 最后复核:2026-06-27T00:22:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

jordann6 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T00:22:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Terraform-managed serverless inventory management API (API Gateway v2 + Lambda + DynamoDB) where each inventory item can be analyzed in real time. Lambda calls Claude 3.5 Haiku via Amazon Bedrock to evaluate stock lev…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production-deployed serverless API with 5 routes (CRUD + /analyze endpoint). Bedrock analysis returns structured JSON with status (critical/warning/ok), recommendation, and reorder_quantity fields. Includes Terraform Ia…

模型作用:Claude 3.5 Haiku 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Haiku (Bedrock model ID: anthropic.claude-3-5-haiku-20241022-v1:0) is the foundation model that receives inventory item data (name, SKU, quantity, reorder threshold, unit cost, category) and performs real-tim…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (0 stars) but well-architected with IaC, CI/CD, encryption, and least-privilege IAM. Model ID explicitly documented in Terraform config.

原始记录:AI-powered inventory reorder recommendations using Claude 3.5 Haiku on AWS Bedrock

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

vijayrajeshr 使用 Claude 3.5 Haiku 处理研究分析和报告生成

vijayrajeshr · Claude 3.5 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3.5 Haiku 来源平台:github 最后复核:2026-06-27T00:22:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

vijayrajeshr 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T00:22:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SentinelOps-AI is an autonomous operations agent built on AWS that continuously ingests infrastructure logs (CloudWatch, S3, Kinesis), uses Claude 3.5 Haiku to understand anomalies and diagnose root causes, then trigger…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production-grade autonomous DevOps pipeline with log ingestion, AI-driven anomaly detection and root cause analysis, and automated remediation against live AWS infrastructure. Architecture includes agent core, MCP too…

模型作用:Claude 3.5 Haiku 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Haiku serves as the reasoning engine in the AI agent core, analyzing infrastructure logs for patterns (errors, latency spikes, service crashes), performing root cause reasoning with infrastructure context, an…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (0 stars). README is comprehensive but repo code structure suggests it's a demo/proof-of-concept. Described as 'production-grade' but actual deployment status unknown.

原始记录:Autonomous self-healing DevOps pipeline powered by Claude 3.5 Haiku for AWS log analysis

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

abhinavchadaga 使用 Claude 3.5 Haiku 处理智能体流程编排

abhinavchadaga · Claude 3.5 Haiku

A
厂商:Anthropic / Claude 模型:Claude 3.5 Haiku 来源平台:github 最后复核:2026-06-27T00:22:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

abhinavchadaga 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T00:22:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A GitHub Action that automatically generates comprehensive pull request descriptions by analyzing PR code changes, commit messages, and metadata using Claude 3.5 Haiku. Includes smart filtering (ignore patterns for gene…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A publishable GitHub Action (v1) with CI/CD badges (Super-Linter, CodeQL, CI, check-dist), structured output format (Summary + Changes Made), configurable ignore patterns, and token caching to reduce API costs. Generate…

模型作用:Claude 3.5 Haiku 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Claude 3.5 Haiku analyzes the complete PR code diff, commit messages, and metadata to generate structured, professional pull request descriptions. Its low cost (~micro-cost per PR) and speed make it practical for contin…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (1 star) but well-built with proper CI/CD, linting, CodeQL security scanning, and test coverage. Publishable as a reusable GitHub Action.

原始记录:GitHub Action for automated PR description generation using Claude 3.5 Haiku

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

web-infra-dev / Midscene.js 使用 Seed1.5 VL 处理代码审查和测试生成

web-infra-dev / Midscene.js · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:23:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

web-infra-dev / Midscene.js 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:23:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Midscene's documentation update introduced configuration for Doubao-1.5-thinking-vision-pro from Volcano Engine in an open-source UI automation framework that writes tests in natural language and drives interfaces throu…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The PR summary states that documentation was updated to introduce and configure Doubao-1.5-thinking-vision-pro in English and Chinese model-provider guides; the Midscene artifact is a public vision-driven UI testing and…

模型作用:Seed1.5 VL 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL acts as the multimodal model behind Midscene-style visual UI interpretation, allowing natural-language test steps to be grounded in screenshots and page state.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a documentation/configuration PR mirrored through a public fork rather than an upstream customer story; it is included because it binds the model to a concrete open-source UI automation product artifact.

原始记录:Midscene documents Doubao-1.5-thinking-vision-pro for vision-driven UI automation

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LobeHub (lobehub/lobehub) 使用 Seed-2.1-Pro-Preview 处理智能体流程编排

LobeHub (lobehub/lobehub) · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:github 最后复核:2026-06-26T16:37:40.925741+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LobeHub (lobehub/lobehub) 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-26T16:37:40.925741+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LobeHub, one of the most popular open-source AI chat interfaces (50K+ GitHub stars), added Doubao-Seed-2.1 models to its VolcEngine provider styling and model registry. This enables all LobeHub users to select and inter…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:LobeHub merged styling and model registry changes for Doubao-Seed-2.1 models, making them available in the model selection dropdown for all VolcEngine provider users. The models appear with proper naming, token limits, …

模型作用:Seed-2.1-Pro-Preview 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Seed 2.1 Pro serves as a selectable LLM backend in LobeHub's multi-provider architecture, providing chat, coding, and agent capabilities to the platform's large user base across web and desktop interfaces.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:LobeHub is a third-party open-source project. The integration is model-registration-only (users need their own VolcEngine API key to use the model).

原始记录:LobeHub open-source AI chat UI adds Doubao-Seed-2.1 model support and styling

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

labring (FastGPT) 使用 Seed-2.1-Pro-Preview 处理知识检索和问答

labring (FastGPT) · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:github 最后复核:2026-06-26T16:37:40.925751+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

labring (FastGPT) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-26T16:37:40.925751+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:FastGPT, an open-source knowledge base and RAG (Retrieval-Augmented Generation) platform, added Doubao Seed 2.1 as a public model preset in its plugin system. The PR aligns the model's context/output limits and modality…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:FastGPT users can now select Doubao-Seed-2.1 from the model preset list for RAG-powered knowledge base Q&A, document summarization, and conversational AI workflows. The preset includes correct context window and token l…

模型作用:Seed-2.1-Pro-Preview 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Seed 2.1 serves as the LLM backend for FastGPT's RAG pipeline, handling context-grounded question answering, document retrieval ranking, and response generation in enterprise knowledge base scenarios.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:FastGPT is a third-party open-source project (labring/Snap-FastGPT). The integration is model-preset-only in the plugin system.

原始记录:FastGPT open-source knowledge base system adds Doubao-Seed-2.1 model preset

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

鲲鹏Talk (Bilibili tech creator) 使用 Seed-2.1-Pro-Preview 处理代码审查和测试生成

鲲鹏Talk (Bilibili tech creator) · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:bilibili 最后复核:2026-06-26T16:37:40.925755+00:00 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

鲲鹏Talk (Bilibili tech creator) 公开的代码审查与测试案例,来源为 bilibili,复核于 2026-06-26T16:37:40.925755+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Developer 鲲鹏Talk integrated Doubao-Seed-2.1-Pro into both ChatGLM and TRAE IDE to test real-world coding performance. The developer used Seed 2.1 Pro via TRAE to build a complete Xiaomi car official website frontend page, testing code organization, visual rendering, and development speed. The API throughput was measured at approximately 58 tokens/second.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A complete Xiaomi car-style website frontend page was generated using Seed 2.1 Pro through TRAE IDE. The developer reported high code quality, good code organization, and excellent visual output. The model completed a c…

模型作用:Seed-2.1-Pro-Preview 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Seed 2.1 Pro provided the code generation capability for the full-stack frontend development task. The model demonstrated strong coding performance (58 tok/s throughput) and was able to produce production-quality code f…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an individual developer's evaluation/demo, not an enterprise deployment. The Bilibili video has 4,855 views and 49 likes, indicating moderate community engagement.

原始记录:Developer uses Seed 2.1 Pro via TRAE IDE to build Xiaomi car website frontend, achieving 58 tok/s throughput

已有真实案例 代码审查与测试bilibiliA 类可核验real_case auto_approved 进入模型卡精选

Shopify (Shopify/reasonableai) 使用 DeepSeek LLM 67B Chat 处理真实任务执行

Shopify (Shopify/reasonableai) · DeepSeek LLM 67B Chat

A
厂商:DeepSeek 模型:DeepSeek LLM 67B Chat 来源平台:github 最后复核:2026-06-27T05:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Shopify (Shopify/reasonableai) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T05:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Shopify's ReasonableAI team benchmarked deepseek-llm:67b-chat (via Ollama) alongside mistral:latest, mixtral:latest, and llama2:13b on query classification prompts for customer support routing. Each model was run 5 time…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:All four models returned the same and consistent classification results across 5 runs with 0.0 temperature. Mistral achieved the best overall performance on the classification task, but DeepSeek LLM 67B Chat demonstrate…

模型作用:DeepSeek LLM 67B Chat 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek LLM 67B Chat provided deterministic, consistent query classification outputs across multiple runs, validating its suitability as a candidate for production customer support routing at Shopify. The evaluation sh…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Benchmarking-only (not confirmed production deployment); evaluation was internal to Shopify ReasonableAI team; specific classification accuracy scores not publicly disclosed in the issue.

原始记录:Shopify ReasonableAI benchmarks DeepSeek LLM 67B Chat for query classification in customer support

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

the-crypt-keeper (Can AI Code proje… 使用 DeepSeek LLM 67B Base 处理真实任务执行

the-crypt-keeper (Can AI Code project) · DeepSeek LLM 67B Base

A
厂商:DeepSeek 模型:DeepSeek LLM 67B Base 来源平台:github 最后复核:2026-06-27T00:35:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

the-crypt-keeper (Can AI Code project) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T00:35:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Can AI Code project (600+ stars), a self-evaluating interview framework for AI coders, evaluated deepseek-ai/deepseek-llm-67b-base for code generation capabilities across multiple programming languages and difficult…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The evaluation was completed (issue #119 closed), producing code generation benchmark results for deepseek-llm-67b-base. The model was evaluated alongside other LLMs including DeepSeek Coder variants, CodeLlama, and Mis…

模型作用:DeepSeek LLM 67B Base 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek LLM 67B Base was the base model under evaluation for its code generation capabilities, tested through the project's automated self-evaluation interview system that scores models on coding tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The evaluation results are stored in ndjson files in the results directory. The model was the subject of evaluation rather than being deployed in production.

原始记录:Can AI Code: DeepSeek LLM 67B Base Code Generation Evaluation

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ThingsIBuild (TheX-23) 使用 DeepSeek LLM 67B 处理研究分析和报告生成

ThingsIBuild (TheX-23) · DeepSeek LLM 67B

A
厂商:DeepSeek 模型:DeepSeek LLM 67B 来源平台:github 最后复核:2026-06-27T00:35:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ThingsIBuild (TheX-23) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T00:35:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Multilingual AI Assistant project integrated DeepSeek LLM 67B alongside Mistral, DeepSeek Coder 33B, GPT-o1, and Mixtral 8X7B as one of its core language models for multilingual text processing across 21+ languages,…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working multilingual AI assistant application supporting 21+ languages with multiple LLM backend options, including DeepSeek LLM 67B for general-purpose multilingual text understanding and generation. The project incl…

模型作用:DeepSeek LLM 67B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek LLM 67B served as one of several LLM backends for the assistant's multilingual text processing capabilities, leveraging the model's strong Chinese and English bilingual training for cross-language tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small open-source project (1 star, 2 forks). The model was added as an option but the project's overall adoption is minimal.

原始记录:Multilingual AI Assistant with DeepSeek LLM 67B for Cross-Language Support

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

IntelliBar (intellibar.app) 使用 DeepSeek R1 (Jan) 处理知识检索和问答

IntelliBar (intellibar.app) · DeepSeek R1 (Jan)

A
厂商:DeepSeek 模型:DeepSeek R1 (Jan) 来源平台:official_web 最后复核:2026-06-27T00:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

IntelliBar (intellibar.app) 公开的知识库与检索问答案例,来源为 官方页面,复核于 2026-06-27T00:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:IntelliBar is a Mac desktop assistant product that specifically highlights DeepSeek R1 as a supported advanced model for context-aware tasks such as editing emails within the Mail app, summarizing articles in the browse…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A commercially available macOS desktop AI assistant that integrates DeepSeek R1 with system-level accessibility APIs to provide inline AI assistance across any Mac application, including email editing, article summariza…

模型作用:DeepSeek R1 (Jan) 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek R1 provides the reasoning backbone for understanding context from any Mac application and generating appropriate responses. Its reasoning capabilities enable the assistant to handle complex, multi-step user req…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Product page describes DeepSeek R1 support; exact implementation details not publicly documented in code.

原始记录:IntelliBar: Mac AI Assistant Leveraging DeepSeek R1 for App-Level Intelligence

已有真实案例 知识库与检索问答官方页面A 类可核验real_case auto_approved 进入模型卡精选

数字游牧人 (Bilibili UP主/程序员) 使用 GLM-4.7 处理软件工程任务执行

数字游牧人 (Bilibili UP主/程序员) · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:bilibili 最后复核:2026-06-26T17:06:15Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

数字游牧人 (Bilibili UP主/程序员) 公开的代码代理与软件工程案例,来源为 bilibili,复核于 2026-06-26T17:06:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:程序员使用智谱 GLM-4.7 作为 Claude Code 的推理后端,完成实际编程开发任务。视频展示了如何将 GLM-4.7 接入 Claude Code 工具链,利用其编程能力完成代码编写、调试和项目开发。该视频获得 28 万播放和 5404 点赞,证明了大量开发者认可该用法。

公开产物:成功将 GLM-4.7 接入 Claude Code 作为编程后端,完成实际代码开发任务。视频展示了完整的接入配置流程和实际编程效果,证明 GLM-4.7 在 agentic coding 场景下的可用性和效果。

模型作用:GLM-4.7 作为 Claude Code 的推理后端,提供编程理解、代码生成和多步推理能力。其增强的编程能力和长上下文支持使其适合作为 coding agent 的底层模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:视频为中文内容,描述基于视频标题和描述的推断;未能提取视频内具体代码输出细节。

原始记录:程序员使用 GLM-4.7 作为 Claude Code 后端进行编程开发

已有真实案例 代码代理与软件工程bilibiliA 类可核验real_case auto_approved 进入模型卡精选

御风大世界 (Bilibili UP主/开发者) 使用 GLM-4.7 处理真实任务执行

御风大世界 (Bilibili UP主/开发者) · GLM-4.7

A
厂商:Z AI / GLM 模型:GLM-4.7 来源平台:bilibili 最后复核:2026-06-26T17:06:15Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

御风大世界 (Bilibili UP主/开发者) 公开的真实任务执行案例,来源为 bilibili,复核于 2026-06-26T17:06:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发者使用 GLM-4.7 完成原本需要数周工作量的前后端全栈开发任务。视频标题明确表示'GLM-4.7 把我几周的工作量一次性梭哈了',展示了前后端实际测试结果和配套提示词,证明模型在真实全栈项目中的高效生产力。

公开产物:使用 GLM-4.7 一次性完成了原本需要数周的前后端开发工作,视频附带了完整的提示词和前后端实测演示,证明 GLM-4.7 在复杂全栈开发场景下的实际产出能力。

模型作用:GLM-4.7 的增强编程能力和多步推理能力使其能够理解复杂项目需求并生成前后端代码,大幅缩短开发周期。模型在代码理解和生成方面的提升是完成此任务的关键。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:视频为中文内容,具体项目细节和代码质量需参考视频实际内容;未能提取视频内具体代码输出。

原始记录:开发者使用 GLM-4.7 完成数周工作量的前后端全栈开发

已有真实案例 真实任务执行bilibiliA 类可核验real_case auto_approved 进入模型卡精选

OpenClaw开源社区 使用 GLM-4.7-Flash 处理智能体流程编排

OpenClaw开源社区 · GLM-4.7-Flash

A
厂商:Z AI / GLM 模型:GLM-4.7-Flash 来源平台:bilibili 最后复核:2026-06-26T17:06:15Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

OpenClaw开源社区 公开的智能体工作流案例,来源为 bilibili,复核于 2026-06-26T17:06:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:OpenClaw 开源社区通过 1Panel + Ollama 一键部署 GLM-4.7-Flash 本地 AI 智能体,实测 Skills 安装、文件整理、Web 搜索等功能。视频展示了完整的本地部署流程,验证 GLM-4.7-Flash 作为轻量 MoE 模型在本地运行时不占用大量资源,真正实现 Token 自由。视频获得 7.2 万播放和 937 点赞。

公开产物:成功在本地环境部署 GLM-4.7-Flash AI 智能体,实测验证了 Skills 安装、文件整理、Web 搜索等功能的可用性。模型在本地运行轻量不占资源,无需支付 API 费用。

模型作用:GLM-4.7-Flash 作为 30B 级 SOTA MoE 模型,其轻量级架构使其适合本地部署,在保持较高性能的同时显著降低资源消耗。模型的中文理解能力使其在本地智能体场景中表现优异。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:视频为中文内容,具体功能测试细节需参考视频实际内容;OpenClaw 为开源项目而非大型企业。

原始记录:OpenClaw 社区使用 GLM-4.7-Flash 本地部署 AI 智能体实现 Token 自由

已有真实案例 智能体工作流bilibiliA 类可核验real_case auto_approved 进入模型卡精选

gameworkerkim 使用 DeepSeek V3 (Dec) 处理研究分析和报告生成

gameworkerkim · DeepSeek V3 (Dec)

A
厂商:DeepSeek 模型:DeepSeek V3 (Dec) 来源平台:github 最后复核:2026-06-26T17:19:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

gameworkerkim 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T17:19:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Built LAON VaultGuard, a git secret detection tool that implements a two-stage pipeline: first a regex-based 'git grep' keyword filter (60+ patterns) for speed, then multi-LLM context analysis (Claude, DeepSeek, GPT, Ol…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Production security tool available as npm package (npx laon-vaultguard scan .), VS Code extension, and Docker container with web dashboard. Uses majority voting across multiple LLMs including DeepSeek V3 to minimize fal…

模型作用:DeepSeek V3 (Dec) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 is one of four LLMs in the multi-LLM context analysis stage, contributing context-aware assessment of whether detected secrets are genuine credentials or false positives. Its role in the ensemble provides co…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Single developer project. DeepSeek V3 is one of four models in ensemble. Primarily useful for Korean-market and global software development teams.

原始记录:LAON VaultGuard — Multi-LLM Cross-Validated Secret Detection

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

disler 使用 DeepSeek V3 (Dec) 处理软件工程任务执行

disler · DeepSeek V3 (Dec)

A
厂商:DeepSeek 模型:DeepSeek V3 (Dec) 来源平台:github 最后复核:2026-06-26T17:19:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

disler 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-26T17:19:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Created an always-on AI assistant pattern specifically powered by DeepSeek V3, combined with RealtimeSTT for speech-to-text and Typer for CLI interaction, designed for software engineering workflows. The system provides…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An open-source pattern/template for building always-on AI assistants with DeepSeek V3 as the core LLM, featuring real-time speech recognition (RealtimeSTT) and a clean CLI interface (Typer). The project demonstrates a p…

模型作用:DeepSeek V3 (Dec) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 serves as the primary reasoning and code generation engine in this always-on assistant architecture. The project specifically chose DeepSeek V3 over alternatives for its combination of coding capability and …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Single developer project (disler). The 'always-on' pattern is a reference architecture, not a production deployment at scale. DeepSeek V3 is the sole LLM in this implementation (no ensemble fallback).

原始记录:Always-On AI Engineering Assistant — DeepSeek-V3 Powered Voice-Activated Development Tool

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CyberStrikeus 使用 DeepSeek V3 (Dec) 处理代码审查和测试生成

CyberStrikeus · DeepSeek V3 (Dec)

A
厂商:DeepSeek 模型:DeepSeek V3 (Dec) 来源平台:github 最后复核:2026-06-26T17:19:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CyberStrikeus 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-26T17:19:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Integrated DeepSeek V3 as a supported LLM backend for CyberStrike, an AI-powered offensive security agent with 7,300+ actionable security skills based on MITRE ATT&CK. DeepSeek V3 is positioned as a cost-effective alter…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An open-source AI pentesting agent with 7,300+ security skills supporting multiple LLM backends including DeepSeek V3. The platform enables autonomous penetration testing workflows where DeepSeek V3 provides the reasoni…

模型作用:DeepSeek V3 (Dec) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 serves as one of the primary LLM backends for CyberStrike's autonomous pentesting capabilities, specifically positioned for cost-effective operation. It processes security assessment tasks including vulnerab…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:DeepSeek V3 is one of many supported LLM backends, not the exclusive model. Security tooling use case. 716 stars indicates emerging rather than established adoption.

原始记录:CyberStrike — Autonomous Penetration Testing Agent with DeepSeek V3 Backend

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

continuedev/continue (popular open-… 使用 DeepSeek V3.1 处理软件工程任务执行

continuedev/continue (popular open-source AI coding IDE extension, 20K+ GitHub stars) · DeepSeek V3.1

A
厂商:DeepSeek 模型:DeepSeek V3.1 来源平台:github 最后复核:2026-06-27T01:06:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

continuedev/continue (popular open-source AI coding IDE extension, 20K+ GitHub stars) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T01:06:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developers use DeepSeek V3.1 (model ID: deepseek-v3.1:671b-cloud) as the LLM backend in Continue.dev for AI-assisted code completion, chat, and multi-file editing within VS Code and JetBrains IDEs. 26 GitHub issues refe…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Continue.dev users running DeepSeek V3.1 for real-time code suggestions, inline edits, and conversational coding assistance. Issues include rate limit handling (429 errors), unknown API errors, and model configuration —…

模型作用:DeepSeek V3.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1 serves as the inference backbone for Continue's AI coding features: code completion, chat-based refactoring, multi-file editing, and context-aware suggestions. Its hybrid thinking mode enables both fast co…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Some users report 429 rate limit errors and occasional unknown errors when using DeepSeek V3.1 via Continue, indicating high demand on DeepSeek API infrastructure.

原始记录:Continue.dev IDE Extension: DeepSeek V3.1 as Primary Coding Model Backend

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

cline/cline (popular open-source AI… 使用 DeepSeek V3.1 处理软件工程任务执行

cline/cline (popular open-source AI coding agent for VS Code/JetBrains/CLI, 28K+ GitHub stars) · DeepSeek V3.1

A
厂商:DeepSeek 模型:DeepSeek V3.1 来源平台:github 最后复核:2026-06-27T01:06:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

cline/cline (popular open-source AI coding agent for VS Code/JetBrains/CLI, 28K+ GitHub stars) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T01:06:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developers use DeepSeek V3.1 as the LLM backend in Cline, an autonomous AI coding agent that reads/writes files, runs terminal commands, browses the web, and builds features through natural conversation. Cline explicitl…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Cline users running DeepSeek V3.1 for autonomous software development tasks including feature implementation, bug fixing, code review, and full-stack development. Issue #8365 'DeepSeek V3.2 always put its XML tool calli…

模型作用:DeepSeek V3.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1 powers Cline's autonomous coding loop: file reading/writing, terminal command execution, browser interaction, and multi-file code generation. Its tool-calling capabilities and 128K context window enable Cl…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Users report token consumption issues and tool-calling format differences between V3.1 and V3.2, suggesting V3.1's tool-calling interface has specific patterns that downstream tools must accommodate.

原始记录:Cline AI Coding Agent: DeepSeek V3.1 as Autonomous Coding Backend

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

jiqi136/Ai-Assistant (open-source A… 使用 DeepSeek V3.1 处理软件工程任务执行

jiqi136/Ai-Assistant (open-source AI assistant aggregating top LLMs, listed in DeepSeek official integration repo) · DeepSeek V3.1

A
厂商:DeepSeek 模型:DeepSeek V3.1 来源平台:github 最后复核:2026-06-27T01:06:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

jiqi136/Ai-Assistant (open-source AI assistant aggregating top LLMs, listed in DeepSeek official integration repo) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T01:06:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A real-time web-access AI assistant that uses DeepSeek-V3.1 as its core LLM interface, enabling direct API access without network relay (costs slashed by 90%). Supports image understanding, web browsing, novel writing, …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:DS-AI assistant application running DeepSeek V3.1 for real-time web-access AI capabilities including image understanding, web browsing, novel writing assistance, and AI programming. The project aggregates multiple top L…

模型作用:DeepSeek V3.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1 provides the primary LLM backbone for DS-AI's multi-capability assistant, enabling code generation, image understanding, web content analysis, and creative writing. Its competitive pricing makes it ideal f…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (26 stars) but verified in DeepSeek's official integration list. The 'costs slashed by 90%' claim refers to using DeepSeek API directly versus other providers.

原始记录:DS-AI Real-time Web-Access AI Assistant: DeepSeek V3.1 as Core Interface

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

BBuf/KDA-Pilot (kernel development … 使用 DeepSeek V3.1 处理研究分析和报告生成

BBuf/KDA-Pilot (kernel development automation pilot, 6.8K+ GitHub stars) · DeepSeek V3.1

A
厂商:DeepSeek 模型:DeepSeek V3.1 来源平台:github 最后复核:2026-06-27T01:06:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

BBuf/KDA-Pilot (kernel development automation pilot, 6.8K+ GitHub stars) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T01:06:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:KDA-Pilot uses DeepSeek V3.1 (referenced as deepseek_v31) for profile-filtered CUDA kernel development tasks. The tool automates kernel optimization by leveraging DeepSeek V3.1's reasoning capabilities to analyze profil…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:KDA-Pilot running DeepSeek V3.1 for automated CUDA kernel profiling analysis and optimization suggestions. PR #118 'deepseek_v31 — profile-filtered kernel tasks' adds V3.1-specific kernel task handling, integrating prof…

模型作用:DeepSeek V3.1 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1's thinking/reasoning mode enables KDA-Pilot to analyze complex CUDA kernel profiling data, identify performance bottlenecks, and generate optimized kernel code. The model's ability to process structured pr…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The project uses V3.1 in a highly specialized domain (GPU kernel optimization). Results are specific to CUDA kernel development workflows.

原始记录:KDA-Pilot: DeepSeek V3.1 for Automated CUDA Kernel Profiling and Optimization

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

浙江大学网络空间安全学院 (ZJU AI Safety) + 华为 (… 使用 DeepSeek R1 (Jan) 处理真实任务执行

浙江大学网络空间安全学院 (ZJU AI Safety) + 华为 (Huawei) · DeepSeek R1 (Jan)

A
厂商:DeepSeek 模型:DeepSeek R1 (Jan) 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

浙江大学网络空间安全学院 (ZJU AI Safety) + 华为 (Huawei) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Fine-tune the full DeepSeek R1 (671B) model for enhanced safety and compliance using multi-dimensional safety training corpus, safety-supervised fine-tuning with core safety reasoning pre-alignment, and safety reinforce…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Open-sourced safety-enhanced DeepSeek R1 model weights published on ModelScope (modelscope.cn/models/ZJUAISafety/DeepSeek-R1-Safe). Training pipeline includes safety data generation (jailbreak attack/defense reinforceme…

模型作用:DeepSeek R1 (Jan) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek R1 is the base model being safety-fine-tuned. Its reasoning capabilities (chain-of-thought, self-verification) are preserved while safety behaviors are enhanced. The project demonstrates DeepSeek R1's adaptabil…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Fine-tuned derivative, not vanilla DeepSeek R1. Hardware requirement is very high (64× Ascend 910B). Academic research project with Huawei partnership.

原始记录:DeepSeek-R1-Safe: Safety-Aligned DeepSeek R1 Fine-Tuned on Huawei Ascend 910B

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

paquino11 使用 DeepSeek R1 (Jan) 处理知识检索和问答

paquino11 · DeepSeek R1 (Jan)

A
厂商:DeepSeek 模型:DeepSeek R1 (Jan) 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

paquino11 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a fully local ChatPDF application using RAG (Retrieval-Augmented Generation) with DeepSeek R1 running via Ollama as the reasoning model. Users can upload PDF documents and interact with them through a Streamlit ch…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Open-source Streamlit application (83 GitHub stars) that runs DeepSeek R1 locally via Ollama for PDF question-answering. Features include multi-PDF upload, customizable retrieval (k results, similarity threshold), memor…

模型作用:DeepSeek R1 (Jan) 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek R1 serves as the local LLM backbone for reasoning over retrieved document chunks. Its chain-of-thought reasoning enables multi-step document comprehension and accurate answer generation from PDF content, all ru…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project, primarily a demonstration/tutorial-style application. Requires local GPU for running DeepSeek R1 via Ollama. Lower star count (83) but has complete, working codebase.

原始记录:ChatPDF-RAG: Local PDF Question-Answering with DeepSeek R1 via Ollama

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

JanPlessow 使用 MiMo-V2-Pro 处理智能体流程编排

JanPlessow · MiMo-V2-Pro

A
厂商:Xiaomi / MiMo 模型:MiMo-V2-Pro 来源平台:github 最后复核:2026-06-27T01:35:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

JanPlessow 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T01:35:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:JanPlessow built a mechanical dispatcher agent in OpenClaw that calls sessions_spawn to delegate subtasks. They tested openrouter/xiaomi/mimo-v2-pro (alongside Grok 4 Fast) and discovered the model fills in ALL optional…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Identified that MiMo-V2-Pro populates all optional tool-schema parameters even when instructed not to, leading to wasted tokens and potential harmful parameter combinations. Proposed and received a config-level fix: per…

模型作用:MiMo-V2-Pro 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Pro was used as the primary coding model for the dispatcher agent. Its specific behavior (filling all optional schema parameters) was the empirical basis for a new OpenClaw feature — per-agent tool schema parame…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Model accessed via openrouter/xiaomi/mimo-v2-pro (routed to V2.5-Pro since V2-Pro was deprecated). The behavior observed is from the routed model. Issue is a feature request with working reproduction, not a success stor…

原始记录:OpenClaw user JanPlessow builds mechanical dispatcher agent using MiMo-V2-Pro for tool-call parameter filtering

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

jadesun 使用 DeepSeek V3.1 处理研究分析和报告生成

jadesun · DeepSeek V3.1

A
厂商:DeepSeek 模型:DeepSeek V3.1 来源平台:github 最后复核:2026-06-26T17:53:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

jadesun 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T17:53:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Automated financial news crawling and AI-powered stock investment recommendation. The system crawls news from East Money (东方财富) using Scrapy, then feeds articles to DeepSeek-V3.1 (accessed via Volcano Engine ARK platform) for deep analysis. The model evaluates news sentiment, market impact, and generates stock investment suggestions with confidence scores and risk assessments. The entire pipeline runs 24/7 in Docker.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Hourly AI-generated stock investment recommendations with confidence scores, risk assessments, and sector categorizations displayed via a responsive Flask web dashboard.

模型作用:DeepSeek V3.1 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V3.1 is the core analysis engine — it receives raw financial news text and produces structured investment recommendations including stock picks, confidence levels, risk ratings, and reasoning. The model's hybri…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Model accessed via third-party Volcano Engine ARK platform endpoint (ep-20250827105540-7wzzj), not direct DeepSeek API. Stock recommendations are AI-generated and should not be taken as financial advice. The repo has 21…

原始记录:jadesun/ai_stock: AI-powered stock recommendation system using DeepSeek-V3.1 for financial news analysis

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

1pperalta 使用 DeepSeek V3.1 处理研究分析和报告生成

1pperalta · DeepSeek V3.1

A
厂商:DeepSeek 模型:DeepSeek V3.1 来源平台:github 最后复核:2026-06-26T17:53:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

1pperalta 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T17:53:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Full-stack real-time sports betting odds analysis platform. A Python scraper collects odds from The Odds API across Premier League, La Liga, Serie A, Bundesliga, Ligue 1, and Champions League. A LangGraph AI agent power…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Structured markdown betting analysis with odds comparisons, best value picks, justifications, and supporting data — rendered in a React 18 frontend with TailwindCSS.

模型作用:DeepSeek V3.1 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1 serves as the reasoning backbone of the LangGraph agent. It processes odds data, RAG-retrieved football statistics, and live standings to generate analytical betting recommendations. Budget controls enforc…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Model accessed via OpenRouter API, not direct DeepSeek endpoint. Repo has 3 followers. Last updated 2026-04-16. Betting analysis is AI-generated and not guaranteed.

原始记录:1pperalta/ScrapOddsAPI: LangGraph AI agent for real-time sports betting odds analysis powered by DeepSeek V3.1

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Independent researcher (togethercom… 使用 DeepSeek LLM 67B Chat 处理研究分析和报告生成

Independent researcher (togethercomputer/MoA issue #41) · DeepSeek LLM 67B Chat

A
厂商:DeepSeek 模型:DeepSeek LLM 67B Chat 来源平台:github 最后复核:2026-06-27T05:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Independent researcher (togethercomputer/MoA issue #41) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T05:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A researcher configured a Mixture of Agents (MoA) pipeline using deepseek-llm-67b-chat as one of five intermediate-layer reference models alongside WizardLM-2-8x22B, Mixtral-8x7B-Instruct, Qwen2-72B-Instruct, and Llama-…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The MoA configuration including deepseek-llm-67b-chat as an intermediate model did not outperform the single-model (no MoA) baseline on GSM8K. Both settings achieved equally on math reasoning, suggesting that for this s…

模型作用:DeepSeek LLM 67B Chat 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek LLM 67B Chat contributed as one of five diverse intermediate reasoning models in the MoA ensemble. Its inclusion provided another perspective in the aggregation step, though the experiment showed diminishing re…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Research experiment, not production deployment; MoA is a specific ensemble technique; single model results not individually reported for deepseek-llm-67b-chat in isolation.

原始记录:Researcher uses DeepSeek LLM 67B Chat as intermediate layer in Mixture of Agents for GSM8K math reasoning

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AIR-hl (research group) 使用 DeepSeek V3.1 Terminus 处理多模态内容处理

AIR-hl (research group) · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AIR-hl (research group) 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Training DAPO (Direct Alignment from Preference Optimization) on DeepSeek-V3.1-Terminus with FP8 format for reinforcement learning to improve reasoning capabilities. The training runs on a 24-node cluster with 8 NVIDIA …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Encountered a dimension mismatch error (576 must be a multiple of block_size0: 128) during FP8 weight quantization in the verl training pipeline. The issue was filed on 2026-01-07 and received 14 comments, indicating ac…

模型作用:DeepSeek V3.1 Terminus 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V3.1-Terminus serves as the base model for RL training. Its 671B-parameter MoE architecture with Mixture-of-Experts design enables efficient training at scale. The model's improved language consistency and agen…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The user encountered a technical error during FP8 training, indicating potential compatibility issues with certain quantization formats. The issue is actively being debugged. The model's large size (671B params) require…

原始记录:AIR-hl RL training with DAPO on DeepSeek-V3.1-Terminus using FP8 quantization

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

April-99 (developer) 使用 DeepSeek V3.1 Terminus 处理代码审查和测试生成

April-99 (developer) · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

April-99 (developer) 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Deploying the quantized DeepSeek-V3.1-Terminus-w4a8-mtp-QuaRot model on Huawei Ascend NPUs using vllm-ascend 0.18.0rc1 inference engine. The deployment targets the w4a8 (4-bit weights, 8-bit activations) quantized varia…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Reported a bug where the model intermittently returns empty content in responses when sending specific request payloads. The issue was filed on the vllm-ascend project tracker, indicating active deployment and testing o…

模型作用:DeepSeek V3.1 Terminus 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V3.1-Terminus is the core inference model being deployed. Its w4a8 quantized variant with QuaRot optimization enables efficient deployment on Ascend NPUs while maintaining model quality. The MTP (Mixture-of-Exp…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The deployment encountered intermittent empty response issues, suggesting the quantized model may have edge-case stability problems on Ascend NPU hardware. The vllm-ascend 0.18.0rc1 version is a release candidate, not a…

原始记录:April-99 deploying DeepSeek-V3.1-Terminus on Huawei Ascend NPUs with vllm-ascend

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zhanwuzhijing / QuantumNous new-api 使用 DeepSeek V3.1 Terminus 处理代码审查和测试生成

zhanwuzhijing / QuantumNous new-api · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T09:42:45Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zhanwuzhijing / QuantumNous new-api 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T09:42:45Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A QuantumNous/new-api user configured DeepSeek V3.1 Terminus through two upstream providers, SiliconFlow using deepseek-ai/DeepSeek-V3.1-Terminus and NVIDIA using deepseek-ai/deepseek-v3.1-terminus, then used the new-ap…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public issue documents a reproducible integration artifact with screenshots: new-api treats the upper-case and lower-case provider model identifiers as the same model for duplicate checks, but the bound-channel disp…

模型作用:DeepSeek V3.1 Terminus 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1 Terminus is the concrete model being exposed through the gateway; its provider-specific identifiers drive the channel binding, routing, and model-management behavior that the user is testing in the product.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub issue with user-supplied configuration and screenshots rather than an official customer story; it is still a real integration/debugging use case with exact model IDs, named user, task, result…

原始记录:QuantumNous/new-api 用户配置 DeepSeek V3.1 Terminus 的 SiliconFlow 与 NVIDIA 双渠道并暴露大小写匹配问题

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

rick-stevens-ai 使用 DeepSeek V3.1 Terminus 处理智能体流程编排

rick-stevens-ai · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T06:03:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

rick-stevens-ai 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:03:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:公开仓库提供在 Intel Data Center GPU Max 1550 上用 llama.cpp SYCL 后端和多 GPU layer splitting 运行 DeepSeek-V3.1-Terminus 671B/405GB Q4_K_M 的完整配置。

公开产物:仓库描述给出实际运行结果:8 张 GPU 上实现 OpenAI-compatible API server,并报告 4.1 tok/s generation across 8 GPUs;README 列出 131K context、9 个 GGUF 分片、硬件/软件要求和启动配置。

模型作用:DeepSeek-V3.1-Terminus 是被部署和服务的核心大模型,用于代码生成、推理和 agent 任务;该项目验证了其在 Intel MAX GPU 多卡环境中的推理可行性。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是工程部署/推理基础设施案例,不是终端业务 customer story;但仓库明确绑定 exact model、组织、任务和可访问产物。

原始记录:rick-stevens-ai 在 8 张 Intel Data Center GPU Max 1550 上运行 DeepSeek-V3.1-Terminus

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

sammy22cool 使用 DeepSeek V3.1 Terminus 处理软件工程任务执行

sammy22cool · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T09:38:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

sammy22cool 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a Node.js/Express proxy exposing OpenAI-compatible /v1/models and /v1/chat/completions endpoints while routing incoming chat completion requests to the NVIDIA NIM model id deepseek-ai/deepseek-v3.1-terminus.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains a runnable Express service, package.json, health check, model-list endpoint, and chat-completions handler; server.js maps common OpenAI, Claude, Gemini, and DeepSeek model aliases to deeps…

模型作用:DeepSeek V3.1 Terminus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3.1 Terminus is the backend inference model for the proxy; all compatible client requests are normalized and routed to the Terminus model so downstream tools can use it through an OpenAI-style API surface.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub repository and source file, not an official customer story; the implementation is self-published by the repository owner but the exact model id is present in code and the artifact is reachabl…

原始记录:sammy22cool built an OpenAI-compatible proxy that routes requests to DeepSeek V3.1 Terminus via NVIDIA NIM

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

going-doer (Paper2Code project) 使用 DeepSeek-Coder-V2-Lite-Instruct 处理研究分析和报告生成

going-doer (Paper2Code project) · DeepSeek-Coder-V2-Lite-Instruct

A
厂商:DeepSeek 模型:DeepSeek-Coder-V2-Lite-Instruct 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

going-doer (Paper2Code project) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Paper2Code automates code generation from scientific papers in machine learning. The tool takes ML research papers and automatically generates executable code implementations, using DeepSeek-Coder-V2-Lite-Instruct as th…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An open-source tool (4,698+ GitHub stars) that automatically generates ML code from papers, with DeepSeek-Coder-V2-Lite-Instruct as the default backend model for all code generation tasks.

模型作用:DeepSeek-Coder-V2-Lite-Instruct 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-Coder-V2-Lite-Instruct serves as the default and primary code generation model in Paper2Code, chosen for its strong code generation capabilities in a local/inference-efficient 16B-parameter MoE architecture wit…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo stars count from search results; exact default branch may differ from 'main'. Model is explicitly stated as default in README.

原始记录:Paper2Code: Automated ML Paper-to-Code Generation Using DeepSeek-Coder-V2-Lite-Instruct

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

eosphoros-ai / DB-GPT 使用 Qwen Chat 14B 处理代码审查和测试生成

eosphoros-ai / DB-GPT · Qwen Chat 14B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 14B 来源平台:github 最后复核:2026-06-27T07:30:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

eosphoros-ai / DB-GPT 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T07:30:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DB-GPT is an open-source AI-native data application framework for private data, SQL generation, knowledge-base QA, dashboards and business insight workflows; its supported-model documentation explicitly lists qwen-14b-c…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public DB-GPT repository provides a runnable framework for local/private-data chat, Text2SQL, RAG and agent workflows, with Qwen-14B-Chat available as one of the local model backends.

模型作用:Qwen Chat 14B 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-14B-Chat supplies the conversational LLM backend used to interpret user questions, generate SQL or answers, and drive private-data assistant interactions inside DB-GPT deployments.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from project documentation and source repository rather than an end-customer production write-up; it is a concrete public software artifact, not a benchmark-only page.

原始记录:DB-GPT supports Qwen-14B-Chat for private data chat, Text2SQL and knowledge-base QA

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Raullen Chai (Rapid-MLX project) 使用 DeepSeek-Coder-V2-Lite-16B-4bit 处理真实任务执行

Raullen Chai (Rapid-MLX project) · DeepSeek-Coder-V2-Lite-16B-4bit

A
厂商:DeepSeek 模型:DeepSeek-Coder-V2-Lite-16B-4bit 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Raullen Chai (Rapid-MLX project) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Rapid-MLX is a local AI inference engine for Apple Silicon (3,106+ stars) that supports DeepSeek-Coder-V2-Lite-16B in 4-bit quantized form as one of its featured model backends. The engine provides 4.2x faster inference…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A high-performance local inference engine enabling developers to run DeepSeek-Coder-V2-Lite on Apple Silicon hardware with optimized MLX backend, achieving significantly faster inference than Ollama while supporting the…

模型作用:DeepSeek-Coder-V2-Lite-16B-4bit 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-Coder-V2-Lite-16B-4bit is explicitly listed in Rapid-MLX's supported model table as one of the featured models, paired with DeepSeek-R1 and V4 Flash. The model's efficient MoE architecture (2.4B active paramete…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Rapid-MLX supports many models; DCV2-Lite is one of several featured. Star count from search results. The model is 4-bit quantized variant.

原始记录:Rapid-MLX: Fastest Apple Silicon Local AI Engine with DeepSeek-Coder-V2-Lite Backend

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

XAI-liacs / Leiden Institute of Adv… 使用 DeepSeek-Coder-V2-16B 处理真实任务执行

XAI-liacs / Leiden Institute of Advanced Computer Science · DeepSeek-Coder-V2-16B

A
厂商:DeepSeek 模型:DeepSeek-Coder-V2-16B 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

XAI-liacs / Leiden Institute of Advanced Computer Science 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:BLADE (Benchmark suite for LLM-driven Automated Design and Evolution) from Leiden University's LIACS institute uses DeepSeek-Coder-V2:16b as the LLM backend for automated algorithm design and evolution. The system evolv…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An academic benchmark suite (42 stars) and PyPI package (iohblade) that demonstrates DeepSeek-Coder-V2:16b effectively generating and evolving optimization algorithm code through LLaMEA (LLM-driven algorithm evolution),…

模型作用:DeepSeek-Coder-V2-16B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-Coder-V2:16b serves as the code generation engine in BLADE's evolutionary loop, responsible for generating new algorithm variants, modifying existing code, and producing optimized metaheuristic implementations.…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:BLADE is primarily a benchmark suite, but DCV2 is used as the actual code generation backend in the demonstrated workflows. The model is optional (other LLMs supported) but is explicitly shown in the README example code.

原始记录:BLADE: LLM-Driven Automated Algorithm Design Using DeepSeek-Coder-V2 as Code Evolution Engine

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

PR0F3S50R 使用 DeepSeek-Code-V2-Lite-Instruct 处理研究分析和报告生成

PR0F3S50R · DeepSeek-Code-V2-Lite-Instruct

A
厂商:DeepSeek 模型:DeepSeek-Code-V2-Lite-Instruct 来源平台:github 最后复核:2026-06-27T01:40:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

PR0F3S50R 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T01:40:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A real-time network traffic and system log analysis pipeline that uses DeepSeek-Code-V2-Lite-Instruct as the core classification model. The system sniffs network packets via Scapy, monitors log files via Watchdog, and f…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Docker-ready security analysis pipeline that classifies network traffic and system logs as malicious or benign in real-time, using DeepSeek-Code-V2-Lite-Instruct for the core classification inference with few-shot con…

模型作用:DeepSeek-Code-V2-Lite-Instruct 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-Code-V2-Lite-Instruct provides the core LLM inference for the security classification pipeline. The model processes network packet data and log entries, applying few-shot learning with a dynamic context window …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (1 star) but genuine deployment. Uses DeepSeek-Code-V2-Lite-Instruct (not the full 236B model). Model served via LM Studio, not direct API.

原始记录:DeepSeek Morpheus Analyzer: Real-Time Network Security Classification with DeepSeek-Code-V2-Lite

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

tanishq-ctrl 使用 DeepSeek V3 (Dec) 处理研究分析和报告生成

tanishq-ctrl · DeepSeek V3 (Dec)

A
厂商:DeepSeek 模型:DeepSeek V3 (Dec) 来源平台:github 最后复核:2026-06-26T18:18:26.239049+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tanishq-ctrl 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T18:18:26.239049+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a modern healthcare consultation chatbot that provides intelligent health-related assistance. The system includes symptom checking, health tips/FAQs, provider matching, medication tracking, and multilingual suppor…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Full-stack healthcare chatbot with symptom checker, health tips, provider directory, medication reminders, and multilingual support. Built with React/TypeScript frontend, Supabase backend, and DeepSeek V3 API for AI rea…

模型作用:DeepSeek V3 (Dec) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 API provides all AI processing capabilities — symptom analysis, health information generation, conversational health assistance, and medical reasoning. Listed as '🧠 DeepSeek V3 API: Advanced AI processing' …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project (low stars). No evidence of regulatory compliance or medical disclaimers. Healthcare use of LLMs carries inherent risk.

原始记录:SympCheck Healthcare Chatbot Using DeepSeek V3 API

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LMSYS Org / SGLang Team 使用 DeepSeek R1 (Jan) 处理真实任务执行

LMSYS Org / SGLang Team · DeepSeek R1 (Jan)

A
厂商:DeepSeek 模型:DeepSeek R1 (Jan) 来源平台:official_web 最后复核:2026-06-27T02:15:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LMSYS Org / SGLang Team 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T02:15:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The SGLang team at LMSYS deployed and optimized DeepSeek R1 (and V3) inference on NVIDIA GB200 NVL72 hardware, implementing prefill-decode disaggregation with FP8 attention, NVFP4 MoE quantization, and large-scale exper…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Achieved 26,156 input tokens/sec and 13,386 output tokens/sec per GPU on DeepSeek R1 for 2000-token input sequences — a 3.8x prefill and 4.8x decode throughput improvement over H100 baseline. With BF16 attention and FP8…

模型作用:DeepSeek R1 (Jan) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek R1's Mixture-of-Experts architecture with fine-grained expert parallelism was the central target model; the 671B-parameter MoE design enabled the massive throughput gains on GB200 NVL72 through large-scale expe…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Blog post describes benchmark/optimization results from a research team rather than end-user production deployment; however LMSYS is a major systems research organization operating real GPU clusters, and the results are…

原始记录:LMSYS/SGLang Team Deploys DeepSeek R1 on NVIDIA GB200 NVL72 with 3.8x/4.8x Throughput Gains

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

TauricResearch 使用 DeepSeek R1 (Jan) 处理研究分析和报告生成

TauricResearch · DeepSeek R1 (Jan)

A
厂商:DeepSeek 模型:DeepSeek R1 (Jan) 来源平台:github 最后复核:2026-06-27T02:15:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

TauricResearch 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T02:15:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:TauricResearch built TradingAgents, an open-source multi-agent LLM framework for financial trading, integrating DeepSeek R1 as a reasoning backbone for market analysis, sentiment analysis, and trading signal generation.…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Open-source framework with 88,000+ GitHub stars supporting DeepSeek R1 for multi-agent financial trading workflows including Research Manager, Trader, and Portfolio Manager agents. Published research paper (arXiv:2412.2…

模型作用:DeepSeek R1 (Jan) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek R1 provides chain-of-thought reasoning for the framework's core agents — Research Manager (financial analysis), Trader (signal generation), and Portfolio Manager (risk assessment) — leveraging its strong reason…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Framework supports multiple LLM providers; DeepSeek R1 is one of several options rather than the sole model. User must configure DeepSeek R1 explicitly. Community actively requests specific DeepSeek R1 versions (issue #…

原始记录:TradingAgents: 88K-Star Multi-Agent Financial Trading Framework with DeepSeek R1 Reasoning Backbone

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Juexiao Zhou, Bin Zhang, Xiuying Ch… 使用 DeepSeek LLM 67B (V1) 处理研究分析和报告生成

Juexiao Zhou, Bin Zhang, Xiuying Chen et al. (King Abdullah University of Science and Technology / Southern University of Science and Technology) · DeepSeek LLM 67B (V1)

A
厂商:DeepSeek 模型:DeepSeek LLM 67B (V1) 来源平台:github 最后复核:2026-06-26T18:17:12.971643+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Juexiao Zhou, Bin Zhang, Xiuying Chen et al. (King Abdullah University of Science and Technology / Southern University of Science and Technology) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T18:17:12.971643+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AutoBA is an autonomous AI agent for conventional multi-omic bioinformatics analysis (WGS, RNA-seq, scRNA-seq, ChIP-seq, spatial transcriptomics). DeepSeek LLM 67B Chat is one of the explicitly supported LLM backends (c…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The agent generates complete, step-by-step bioinformatics analysis plans and executes them locally. Validated by expert bioinformaticians across whole genome sequencing, RNA-seq, single-cell RNA-seq, ChIP-seq, and spati…

模型作用:DeepSeek LLM 67B (V1) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek LLM 67B Chat serves as the reasoning backbone for the AutoBA agent, enabling it to understand bioinformatics data variations, autonomously design analysis pipelines, and generate executable step-by-step plans. …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:AutoBA supports multiple LLM backends (GPT-4, deepseek-coder, etc.); deepseek-llm-67b-chat is one of several supported options rather than the sole model. No published ablation comparing deepseek-llm-67b-chat against ot…

原始记录:AutoBA: Automated Multi-Omic Bioinformatics Analysis Agent Powered by DeepSeek LLM 67B Chat

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SGLang Project (sgl-project/sglang) 使用 DeepSeek-V2 处理真实任务执行

SGLang Project (sgl-project/sglang) · DeepSeek-V2

A
厂商:DeepSeek 模型:DeepSeek-V2 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SGLang Project (sgl-project/sglang) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Fix correctness issues and optimize performance for DeepSeek-V2's MoE expert routing and MLA attention in SGLang when processing padded forward batches, ensuring accurate model outputs during production inference.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Merged PR fixing MoE/MLA correctness bugs in SGLang's DeepSeek-V2 implementation, specifically addressing issues with padded batch handling that caused incorrect expert routing and attention computation. The fix ensures…

模型作用:DeepSeek-V2 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2's complex MoE architecture with 160 routed experts and MLA attention exposed edge cases in batched inference that required careful handling of padding tokens. The model's architectural innovations drove impr…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a bugfix PR for a specific edge case in batched inference. The fix was needed because V2's architecture is more complex than standard transformer models.

原始记录:SGLang fixes DeepSeek-V2 MoE/MLA correctness and performance with padded forward batches

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

tambehimanshu 使用 DeepSeek-V2 处理代码审查和测试生成

tambehimanshu · DeepSeek-V2

A
厂商:DeepSeek 模型:DeepSeek-V2 来源平台:github 最后复核:2026-06-27T02:10:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tambehimanshu 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T02:10:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developed an intelligent question-answering system that ingests long documents (PDFs, reports, articles) and answers user queries using DeepSeek-V2 as the core language model for understanding context and generating acc…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working QA application that processes long-form documents of various formats, chunks and indexes content, and uses DeepSeek-V2 to generate contextually accurate answers to user questions about the document corpus

模型作用:DeepSeek-V2 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2's long context window and strong language comprehension capabilities enable the system to maintain coherence when answering questions about lengthy documents, processing up to 128K tokens of context in a sin…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small personal project (3 stars). Repo description explicitly states 'using DeepSeek-V2'. Could not verify README content due to network limitations.

原始记录:Intelligent long-document QA system powered by DeepSeek-V2

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

qinscode 使用 DeepSeek-V2.5 (Dec) 处理研究分析和报告生成

qinscode · DeepSeek-V2.5 (Dec)

A
厂商:DeepSeek 模型:DeepSeek-V2.5 (Dec) 来源平台:github 最后复核:2026-06-26T18:22:09.049962+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

qinscode 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-26T18:22:09.049962+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Automated scraping of IT job listings from Seek.com.au across multiple Australian regions (Perth, Sydney, Melbourne, etc.), with DeepSeek-V2.5 used as the AI backend for post-processing: extracting tech stacks from job …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A production job scraping pipeline that stores structured job data (title, company, salary, tech stack, region) in PostgreSQL, with AI-extracted fields powered by DeepSeek-V2.5 via SiliconFlow API. The system runs on a …

模型作用:DeepSeek-V2.5 (Dec) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 is configured as the default AI_MODEL (deepseek-ai/DeepSeek-V2.5 via SiliconFlow) for two core AI tasks: (1) tech_stack_analyzer.py extracts technology keywords from raw HTML job descriptions, and (2) sala…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The repo has 41 stars and appears to be a personal project. V2.5 is accessed via SiliconFlow proxy, not directly via DeepSeek API. The .env.example explicitly sets AI_MODEL=deepseek-ai/DeepSeek-V2.5 as the default.

原始记录:SeekSpider: AI-Powered Australian IT Job Scraping with DeepSeek-V2.5

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

1596941391qq 使用 DeepSeek-V2.5 (Dec) 处理真实任务执行

1596941391qq · DeepSeek-V2.5 (Dec)

A
厂商:DeepSeek 模型:DeepSeek-V2.5 (Dec) 来源平台:github 最后复核:2026-06-26T18:22:09.049989+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

1596941391qq 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T18:22:09.049989+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Modified Mem0 framework for emotion companion AI, using DeepSeek-V2.5 as the core LLM to extract and categorize user memories from multi-turn conversations. The system separates temporal memories (e.g., 'went skiing on …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A working memory extraction pipeline that processes batches of conversation records (recommended 10 at a time), uses V2.5 to extract both user and AI memories, and stores them in a vector database with temporal/constant…

模型作用:DeepSeek-V2.5 (Dec) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 is configured as the default LLM (model: deepseek-ai/DeepSeek-V2.5 via SiliconFlow) for the memory extraction pipeline. It processes raw conversation transcripts and extracts structured memory entities. Te…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repo has 0 GitHub stars. This is a fork/modification of Mem0 (mem0ai/mem0). The developer's customizations focus on Chinese LLM compatibility and emotion companion use case. V2.5 is used via SiliconFlow API proxy.

原始记录:Mem1: DeepSeek-V2.5 as LLM Backend for Emotion Companion Memory Extraction

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

XInTheDark 使用 DeepSeek-V2.5 (Dec) 处理真实任务执行

XInTheDark · DeepSeek-V2.5 (Dec)

A
厂商:DeepSeek 模型:DeepSeek-V2.5 (Dec) 来源平台:github 最后复核:2026-06-26T18:22:09.049996+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

XInTheDark 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-26T18:22:09.049996+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Raycast extension that provides free access to multiple LLMs including DeepSeek-V2.5 via DeepInfra. V2.5 is listed as an active, working provider with medium speed and 7.5/10 quality rating. The extension supports cha…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Raycast extension (1073 stars) that routes requests to DeepInfra's free tier for DeepSeek-V2.5 access. The model is marked as active (green badge) with medium response speed. Users can access V2.5 for free through the…

模型作用:DeepSeek-V2.5 (Dec) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 is one of the primary models offered through the extension, listed with active status on DeepInfra. The extension enables zero-cost access to V2.5's general conversation and coding capabilities, making it …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a multi-model access tool, not a V2.5-specific application. V2.5 is accessed through DeepInfra's infrastructure, not DeepSeek's own API. The extension also supports many other models. V2.5's active status was ve…

原始记录:Raycast G4F Extension: Free DeepSeek-V2.5 Access via DeepInfra in Raycast

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenDILab 使用 DeepSeek-V2.5 (Dec) 处理文档理解和结构化处理

OpenDILab · DeepSeek-V2.5 (Dec)

A
厂商:DeepSeek 模型:DeepSeek-V2.5 (Dec) 来源平台:github 最后复核:2026-06-26T18:22:09.049999+00:00 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenDILab 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-26T18:22:09.049999+00:00。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenDILab's CleanS2S project is a high-quality streaming speech-to-speech interactive agent implemented in a single file. DeepSeek-V2.5 is specifically recommended and documented as the default local LLM option for the …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully functional speech-to-speech agent (536 stars) that processes real-time voice input, routes through an LLM for conversation, and generates voice output with prosody transfer. The project supports both API-based a…

模型作用:DeepSeek-V2.5 (Dec) 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 is the specifically recommended and documented local LLM option for the speech agent's conversational intelligence. The README provides explicit instructions for running V2.5 locally as the LLM backbone, i…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:CleanS2S supports multiple LLM backends; V2.5 is the recommended option, not the only one. The project is from OpenDILab (536 stars), a credible AI research organization. V2.5's role is as the conversational reasoning e…

原始记录:CleanS2S: DeepSeek-V2.5 as LLM Backend for Streaming Speech-to-Speech Agent

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

vLLM Project 使用 DeepSeek-V2 处理软件工程任务执行

vLLM Project · DeepSeek-V2

A
厂商:DeepSeek 模型:DeepSeek-V2 来源平台:github 最后复核:2026-06-27T05:44:43Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

vLLM Project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:44:43Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The vLLM project integrated DeepSeek-V2/DeepSeek-V2-Chat into its open-source inference engine so users could run the model through vLLM with tensor parallelism and generate responses from prompts.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The pull request was merged on 2024-06-28 and includes a tested example loading model="deepseek-ai/DeepSeek-V2-Chat" with tensor_parallel_size=8, generating an answer to the prompt "The future of AI is?".

模型作用:DeepSeek-V2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2 is the model being served and used for text generation; vLLM's integration makes the model available as a deployable inference target in the vLLM runtime rather than only as a standalone model repository.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an infrastructure integration case, not an end-user business deployment. Evidence is strong for exact model binding and a merged public artifact, but the use case is model serving/testing rather than a customer …

原始记录:vLLM merged DeepSeek-V2 support for high-throughput open-source inference

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

llama.cpp community (ggml-org/llama… 使用 DeepSeek-V2 处理代码审查和测试生成

llama.cpp community (ggml-org/llama.cpp) · DeepSeek-V2

A
厂商:DeepSeek 模型:DeepSeek-V2 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

llama.cpp community (ggml-org/llama.cpp) 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Implement optimized Multi-head Latent Attention (MLA) inference kernels in llama.cpp for DeepSeek-V2/V3, enabling efficient local inference of the 236B MoE model on consumer-grade hardware with quantized GGUF weights.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Merged PR implementing tensor-core optimized MLA decode kernels for DeepSeek-V2/V3 in llama.cpp. The implementation supports GGUF-quantized model variants (Q4_K_M, Q5_K_M, etc.), enabling users to run DeepSeek-V2 locall…

模型作用:DeepSeek-V2 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2's MLA architecture required specialized attention kernels different from standard MHA/GQA. The latent KV compression design enabled significantly lower memory footprint per token, making local inference of t…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:llama.cpp is primarily a local inference engine; this does not represent production serving at scale. Quantized models may have quality degradation compared to full-precision inference.

原始记录:llama.cpp implements optimized MLA inference for DeepSeek-V2/V3 on consumer hardware

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Shaunwei (RealChar open-source proj… 使用 Claude 2.0 处理多模态内容处理

Shaunwei (RealChar open-source project) · Claude 2.0

A
厂商:Anthropic / Claude 模型:Claude 2.0 来源平台:github 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Shaunwei (RealChar open-source project) 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RealChar is an open-source AI character companion platform (6K+ GitHub stars) that uses Claude 2 as one of its primary LLM backends to power real-time conversational AI characters. Users create and customize AI characte…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:RealChar shipped a production application supporting Claude 2 alongside OpenAI, Anyscale Llama2, and other LLMs. The platform provides real-time voice and text chat with AI characters, integrating Claude 2 for high-qual…

模型作用:Claude 2.0 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2 was chosen as a primary LLM backend for RealChar due to its natural conversation quality and safety features. The model powered the character dialogue engine, enabling realistic, engaging, and contextually appr…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Claude 2 is explicitly listed as a supported LLM in the project README. The project has 6K+ GitHub stars and active development. This is a real production use of Claude 2 API, not just a wrapper.

原始记录:RealChar uses Claude 2 as backend LLM for real-time AI character conversations

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LangChain 使用 Claude 2.0 处理知识检索和问答

LangChain · Claude 2.0

A
厂商:Anthropic / Claude 模型:Claude 2.0 来源平台:github 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LangChain 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LangChain's OpenGPTs project is an open-source platform for building custom GPT-like assistants (6K+ GitHub stars). It supports Claude 2 as one of its primary LLM backends, allowing users to create personalized AI assis…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:OpenGPTs shipped as a production-ready open-source platform where users can build and deploy custom AI assistants using Claude 2 as the underlying model. The platform supports RAG pipelines, tool use, and multi-turn con…

模型作用:Claude 2.0 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2 was integrated as a first-class LLM option in OpenGPTs, enabling developers to build custom assistants with Claude 2's reasoning capabilities, 100K context window, and natural language understanding. Users set …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Claude 2 is explicitly supported as a model option in the README. LangChain is a major AI framework organization. This is a production integration, not a tutorial or demo.

原始记录:LangChain OpenGPTs integrates Claude 2 as LLM backend for custom GPT creation

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

smol-ai (GodMode project) 使用 Claude 2.0 处理文档理解和结构化处理

smol-ai (GodMode project) · Claude 2.0

A
厂商:Anthropic / Claude 模型:Claude 2.0 来源平台:github 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

smol-ai (GodMode project) 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:GodMode is an open-source AI chat browser (5K+ GitHub stars) that provides a unified keyboard-shortcut interface to access ChatGPT, Claude 2, Perplexity, Bing and more. Claude 2 was described as 'Excellent, long context…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:GodMode shipped as a production desktop application (Mac, Windows, Linux) that integrates Claude 2's web interface with a custom browser, enabling users to quickly access Claude 2 alongside other AI models via a single …

模型作用:Claude 2.0 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2 was recognized by the GodMode team as offering excellent long-context, multi-document capabilities with fast response times. The model was a core part of GodMode's multi-model strategy, providing a complementar…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Claude 2 is explicitly listed as a supported model in the README with a positive review. The project has 5K+ GitHub stars and active user base. This is a production tool that people used daily.

原始记录:GodMode AI chat browser integrates Claude 2 for fast multi-model access

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

liuqjox (via kvcache-ai/ktransforme… 使用 DeepSeek-V2.5 (Dec) 处理真实任务执行

liuqjox (via kvcache-ai/ktransformers) · DeepSeek-V2.5 (Dec)

A
厂商:DeepSeek 模型:DeepSeek-V2.5 (Dec) 来源平台:github 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

liuqjox (via kvcache-ai/ktransformers) 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:用户 liuqjox 使用 KTransformers 框架在 Windows/WSL 环境下本地部署量化版 DeepSeek-V2.5-1210 (Q2_K_L 量化, 236B 参数, 80.5GB)。硬件配置为 RTX 3090 + 96GB DDR4 3200 + i7-10700K, 环境为 Win10/WSL/Ubuntu/CUDA 12.4/PyTorch 2.6。目标是在不到一万元的消费级硬件上运行 236B 参数大模型进行本地推理。

公开产物:成功实现本地推理,速度稳定在 4-6 tokens/s(初始较慢约 1 tps)。同时验证了更大量化的 DeepSeek-R1 671B (unsloth 2.51 bit, 212GB) 也可运行但速度极慢(0.05 tps)。用户指出 x99 主板 + 256G 内存方案成本不到一万元即可运行 671B 模型。

模型作用:DeepSeek-V2.5-1210 提供了 236B 参数的开放权重模型,其 MoE 架构配合 KTransformers 的 CPU/GPU 异构推理技术,使得在单卡 RTX 3090 + 96GB 内存的消费级硬件上即可实现可用的本地推理(4-6 tps),证明了该模型在资源受限场景下的实用价值。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Issue 状态为 open 但用户已确认部署成功;量化版本 Q2_K_L 可能有精度损失;速度数据基于特定硬件配置,不同配置可能差异较大。

原始记录:KTransformers 用户在消费级硬件上本地部署 DeepSeek-V2.5-1210 236B 实现高效推理

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Alibaba / DingTalk 使用 DeepSeek V3 (Dec) 处理智能体流程编排

Alibaba / DingTalk · DeepSeek V3 (Dec)

A
厂商:DeepSeek 模型:DeepSeek V3 (Dec) 来源平台:github 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Alibaba / DingTalk 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DingTalk (Alibaba's enterprise communication platform with 700M+ registered users) integrated DeepSeek models into its AI Assistant feature, allowing enterprise users to independently select DeepSeek-V3 671B Full-Power …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:DingTalk AI Assistant now offers three DeepSeek model variants to users: DeepSeek-R1 32B Distilled Version, DeepSeek-R1 671B Full-Power Version, and DeepSeek-V3 671B Full-Power Version. The V3 model is specifically list…

模型作用:DeepSeek V3 (Dec) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V3 671B serves as the general-purpose reasoning backbone for DingTalk's AI Assistant, providing enterprise users with high-quality text generation, summarization, and conversational AI capabilities at scale acr…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The DingTalk integration is confirmed in DeepSeek's official awesome-deepseek-integration repository. DingTalk AI Assistant offers multiple model tiers, so DeepSeek V3 is one of several available models, not the sole ba…

原始记录:DingTalk AI Assistant — Enterprise Communication Platform Offering DeepSeek-V3 671B as a Selectable Model

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

rollingfruit (independent developer… 使用 Gemini 1.5 Flash (May) 处理软件工程任务执行

rollingfruit (independent developer, China) · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:GitHub + Bilibili 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

rollingfruit (independent developer, China) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Built a Flask-based real-time visual chat assistant that captures webcam video frames, sends them to Gemini 1.5 Flash for image understanding, and generates spoken responses via ElevenLabs TTS. The system supports free-…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Open-source project (43 stars) with live Bilibili video demo showing real-time visual conversation. The developer reports Gemini 1.5 Flash delivers fast visual response times suitable for low-latency interactive use cas…

模型作用:Gemini 1.5 Flash (May) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Flash is the core vision-language model that processes webcam frames and generates contextual text responses. Its multimodal capability and low latency are explicitly cited as enabling the real-time visual co…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Project is an individual developer demo, not a production deployment. The Bilibili video (BV1Nn4y1o7CK) is the primary artifact demonstrating real usage.

原始记录:Free-Astra: Real-Time Visual AI Conversation via Webcam Using Gemini 1.5 Flash

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ShmuelRonen (open-source contributo… 使用 Gemini 1.5 Flash (May) 处理研究分析和报告生成

ShmuelRonen (open-source contributor) · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:GitHub 最后复核:2026-06-27T02:50:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ShmuelRonen (open-source contributor) 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T02:50:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developed a custom node for the ComfyUI workflow engine that integrates Gemini 1.5 Flash 002 for text generation, image analysis, video processing, and audio transcription within node-based AI pipelines. Supports the fu…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Published ComfyUI custom node (30 stars) enabling AI artists and developers to use Gemini 1.5 Flash in their ComfyUI workflows. The node handles text, image, video, and audio inputs with configurable parameters includin…

模型作用:Gemini 1.5 Flash (May) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Flash 002 is the sole AI model powering all content processing tasks in the node. Its multimodal capabilities and 1M token context window are the key features leveraged for handling diverse media inputs.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Community-contributed node; author later released a 2.0 version for newer Gemini models, indicating 1.5 Flash was the initial production choice.

原始记录:ComfyUI_Gemini_Flash: Custom ComfyUI Node for Multimodal AI Workflows Using Gemini 1.5 Flash 002

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Qodo (formerly CodiumAI) 使用 Gemini 1.5 Pro (May) 处理代码审查和测试生成

Qodo (formerly CodiumAI) · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:qodo.ai 最后复核:2026-06-27T19:00:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Qodo (formerly CodiumAI) 公开的代码审查与测试案例,来源为 qodo.ai,复核于 2026-06-27T19:00:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Qodo integrated Gemini 1.5 Pro into their Qodo Gen v0.12 product (VS Code and JetBrains plugins) for automated code review, test generation, and code understanding tasks across multi-repo codebases.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Developers can now select Gemini 1.5 Pro within Qodo Gen to generate tests, review code, and understand complex codebases. Gemini 1.5 Pro excels at tasks requiring extensive input processing with its 2M token context wi…

模型作用:Gemini 1.5 Pro (May) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Pro's exceptionally large context window (2 million tokens) makes it effective for coding tasks requiring processing extensive input or maintaining coherent understanding across multiple parts of a codebase, …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Qodo blog is the company's own announcement. HN discussion confirms real product integration with user testing feedback.

原始记录:Qodo (formerly CodiumAI) Integrates Gemini 1.5 Pro for Automated Code Review and Test Generation

已有真实案例 代码审查与测试qodo.aiA 类可核验real_case auto_approved 进入模型卡精选

ViaAnthroposBenevolentia 使用 Gemini 2.0 Flash (exp) 处理研究分析和报告生成

ViaAnthroposBenevolentia · Gemini 2.0 Flash (exp)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash (exp) 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ViaAnthroposBenevolentia 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer built a vanilla JavaScript web interface for Gemini 2.0 Flash Multimodal Live API, enabling real-time text chat, audio input/output, webcam video streaming, screen sharing, and function calling — all as a ligh…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully functional open-source web app (388 GitHub stars, live demo on GitHub Pages) that demonstrates Gemini 2.0 Flash's multimodal live capabilities with text, audio, camera, and screen inputs. Simplified from Google'…

模型作用:Gemini 2.0 Flash (exp) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash (exp) is the core model powering all real-time multimodal interactions — text understanding, audio generation, video analysis, and function calling execution in the Live API streaming protocol.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source personal project, not enterprise production deployment. However, the repo is actively maintained (updated 2026-06-21) with substantial community adoption (388 stars).

原始记录:Gemini 2.0 Flash Multimodal Live API Demo — Real-time audio/video/screen interaction client

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Shmuel Ronen 使用 Gemini 2.0 Flash (exp) 处理研究分析和报告生成

Shmuel Ronen · Gemini 2.0 Flash (exp)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash (exp) 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Shmuel Ronen 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer Shmuel Ronen built a ComfyUI custom node integrating Gemini 2.0 Flash for multimodal analysis of text, images, video frames, and audio, plus image generation capabilities via the gemini-2.0-flash-exp-image-gen…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A popular open-source ComfyUI node (335 GitHub stars) enabling AI-powered image generation and multimodal content analysis within Stable Diffusion/ComfyUI workflows. Installable via ComfyUI Manager.

模型作用:Gemini 2.0 Flash (exp) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash (exp) provides both multimodal understanding (text/image/video/audio analysis) and image generation. The model's native multimodal capabilities enable seamless integration into creative AI workflows whe…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source creative tool project. Uses both gemini-2.0-flash-exp for analysis and gemini-2.0-flash-exp-image-generation for image creation. Actively maintained with 335 community stars.

原始记录:ComfyUI-Gemini_Flash_2.0_Exp — Multimodal analysis and image generation node for ComfyUI workflows

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lfglfg11 使用 Gemini 2.0 Flash (exp) 处理研究分析和报告生成

lfglfg11 · Gemini 2.0 Flash (exp)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash (exp) 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

lfglfg11 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer built a Vue.js 3 front-end multimodal assistant using both gemini-2.0-flash-exp and gemini-2.0-flash-exp-image-generation models. Supports intelligent conversation, image recognition, AI image generation, imag…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A fully functional Chinese/English bilingual multimodal assistant app (26 GitHub stars) featuring text chat, image upload/analysis, AI-generated images embedded in conversation, image editing, and responsive design. Dep…

模型作用:Gemini 2.0 Flash (exp) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash (exp) handles all text conversation and image understanding tasks, while the image generation variant powers creative visual output. The model's native multimodal support enables the full chat→image→edi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Personal open-source project by Chinese developer. No backend dependency — all inference goes directly from browser to Google API. Community adoption moderate (26 stars).

原始记录:Gemini 2.0 Multimodal Assistant — Vue.js chat/image recognition/AI image generation app using gemini-2.0-flash-exp

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

daymade 使用 Gemini 2.0 Flash (exp) 处理文档理解和结构化处理

daymade · Gemini 2.0 Flash (exp)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash (exp) 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

daymade 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer daymade published a Python CLI tool on PyPI that uses Gemini 2.0 Flash to generate animated GIFs from text prompts. Features include customizable animation subject/style/frame rate, automatic retry logic for m…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A published PyPI package (gemini-gif) with 17 GitHub stars, providing both CLI and programmatic interfaces for generating animated GIFs. Includes architecture documentation and comprehensive release guide. Produces mult…

模型作用:Gemini 2.0 Flash (exp) 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash (exp) generates individual animation frames from text prompts with sufficient visual consistency to create smooth animated GIFs when composited. The model's image generation capability and speed make it…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Published on PyPI as a distributed package. Personal project inspired by @Miyamura80's gist. Open-source under MIT license. Moderate community adoption.

原始记录:Gemini GIF Generator — Python CLI tool on PyPI for animated GIF generation via Gemini 2.0 Flash

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lesteroliver911 使用 Gemini 2.0 Flash (exp) 处理研究分析和报告生成

lesteroliver911 · Gemini 2.0 Flash (exp)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash (exp) 来源平台:github 最后复核:2026-06-27T02:46:56Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

lesteroliver911 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T02:46:56Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Developer lesteroliver911 built a Python RAG application that uses Gemini 2.0 Flash Exp's vision capabilities to analyze PDF documents. The app extracts content from PDFs, builds a retrieval-augmented generation pipelin…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A functional RAG application (10 GitHub stars) that transforms any PDF into an interactive knowledge base. Leverages Gemini 2.0 Flash Exp's vision-language understanding for processing PDF pages, with RAG implementation…

模型作用:Gemini 2.0 Flash (exp) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash (exp) provides both the vision-language understanding for processing PDF page images and the language model capability for generating context-aware answers to user queries. The model's multimodal input …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Personal project with 10 GitHub stars. Requires Google API key. Uses vision-language capabilities which may have content limits on complex PDF layouts.

原始记录:PDFGenius — RAG application using Gemini 2.0 Flash Exp for interactive PDF knowledge base analysis

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

sinameraji 使用 Kimi K2 处理软件工程任务执行

sinameraji · Kimi K2

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 来源平台:github 最后复核:2026-06-27T06:01:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

sinameraji 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a terminal-native coding agent that uses the Kimi K2 family through Cloudflare Workers AI, with optional AI Gateway routing for observability, caching, per-turn cost logging, and account-owned request telemetry.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository for KimiFlare, described as a terminal coding agent powered by Kimi K2.7 on Cloudflare Workers AI, with documented interactive TUI, plan/edit/auto modes, image input support, shell and file tool…

模型作用:Kimi K2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2 provides the core coding-agent intelligence: reading and editing code, planning multi-step changes, using terminal/file/browser tools, handling screenshots or image prompts, and producing code or command decisio…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub README and repository. The README targets the Kimi K2.7 variant exposed by Cloudflare Workers AI, so this is bound to the Kimi K2 family rather than only the initial K2 release. NPM package p…

原始记录:sinameraji built KimiFlare, a terminal coding agent powered by Kimi K2 on Cloudflare Workers AI

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AliSajid55 / DevLens 使用 Gemini 1.5 Pro (May) 处理知识检索和问答

AliSajid55 / DevLens · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:github 最后复核:2026-06-27T05:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AliSajid55 / DevLens 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T05:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:DevLens 是一个 FastAPI 后端,用户输入任意公开 GitHub 仓库 URL 后,系统索引代码并允许开发者用自然语言询问代码库问题。README 明确标注 LLM 为 Google Gemini 1.5 Pro,并说明 Gemini 负责生成带文件路径、行号和代码片段引用的答案。

公开产物:公开仓库提供可运行的代码库问答 API、Docker Compose 部署方式、SSE 流式回答、Pinecone 命名空间隔离和健康检查端点;示例输出会指出数据库连接等问题对应的具体文件与行号。

模型作用:Gemini 1.5 Pro 作为回答生成模型,在检索到的代码上下文上生成自然语言解释,并把答案约束到真实源代码位置与片段,降低代码库问答中的幻觉风险。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;属于开源产品/项目案例,未独立验证线上生产流量。

原始记录:DevLens 用 Gemini 1.5 Pro 构建 GitHub 代码库问答 RAG API

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chouaibMo / ChatGemini 使用 Gemini 1.5 Pro (May) 处理真实任务执行

chouaibMo / ChatGemini · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:github 最后复核:2026-06-27T05:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chouaibMo / ChatGemini 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T05:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ChatGemini 是一个 Compose Multiplatform 聊天应用,支持 Android、iOS 和桌面端。README 描述其支持文本输入生成、图文多模态输入、多语言、Markdown 和用户自填 API key;代码中的 GeminiService.kt 明确调用 v1beta/models/gemini-1.5-pro:generateContent。

公开产物:公开仓库包含跨平台客户端与截图资源,用户可配置 Gemini API key 后运行聊天功能,在多端进行文本和图文对话。

模型作用:Gemini 1.5 Pro 提供核心 generateContent 能力,负责根据纯文本或图文输入生成聊天回复,是应用文本生成和多模态对话功能的后端模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开代码与 README;项目注明仍在开发中,部分功能可能未完整覆盖所有平台。

原始记录:ChatGemini 用 Gemini 1.5 Pro 实现跨 Android/iOS/Desktop 的多模态聊天应用

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Jatin-Mehra119 / AI-agent-based Dee… 使用 Gemini 1.5 Pro (May) 处理软件工程任务执行

Jatin-Mehra119 / AI-agent-based Deep Research · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:github 最后复核:2026-06-27T05:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Jatin-Mehra119 / AI-agent-based Deep Research 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:该项目是一个基于 LangGraph 的研究自动化系统,用于从公司名称出发并行检索资料、筛选文档、撰写结构化商业情报报告并导出 PDF。README 在技术栈中明确列出 Google Gemini 1.5 Pro API,并说明其用于 AI model interactions、research node 协调和 report generation。

公开产物:公开仓库提供 Streamlit 应用、LangGraph 工作流、Tavily 搜索集成、引用管理和 PDF 导出流程;用户可输入研究目标并得到带引用的结构化研究报告。

模型作用:Gemini 1.5 Pro 承担研究流程中的语言模型交互、信息综合和报告生成,把检索到的网页资料整理成标准章节、引用和最终 PDF 内容。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;在线 Streamlit demo 链接存在跳转行为,因此 artifact 采用可访问的 GitHub 仓库。

原始记录:AI-agent-based Deep Research 用 Gemini 1.5 Pro 自动生成公司研究报告和 PDF

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AntraTripathi74 / Sentiment-Analysis 使用 Gemini 1.5 Pro (May) 处理研究分析和报告生成

AntraTripathi74 / Sentiment-Analysis · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:github 最后复核:2026-06-27T05:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AntraTripathi74 / Sentiment-Analysis 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T05:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:该 Streamlit 应用允许上传销售电话转录 DOCX 文件,使用 Vertex AI GenerativeModel("gemini-1.5-pro") 分析客户情绪、总结关键对话点,并给销售代表提供改进建议。README 明确说明使用 Gemini 1.5 Pro 分析上传的 sales call transcripts。

公开产物:公开仓库包含 app.py 和 README;应用输出客户意向分类(Interested / Not Interested)、要点摘要以及销售代表表现改进建议。

模型作用:Gemini 1.5 Pro 负责读取和理解较长电话转录文本,完成情绪判断、摘要生成和建议生成,是该销售通话分析工作流的核心推理与生成模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开代码与 README;仓库名存在拼写 Anlaysis,但链接可访问。

原始记录:Sentiment-Analysis 用 Gemini 1.5 Pro 分析销售电话 DOCX 转录稿

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lucataco 使用 DeepSeek-V2 处理软件工程任务执行

lucataco · DeepSeek-V2

A
厂商:DeepSeek 模型:DeepSeek-V2 来源平台:github 最后复核:2026-06-27T05:47:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

lucataco 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:47:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:lucataco created a public Cog implementation of deepseek-ai/DeepSeek-V2 so the model can be built and run with Cog/Replicate-style prediction workflows, including a documented `cog predict -i prompt=...` command.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains a working prediction interface (`predict.py`) and Cog build configuration (`cog.yaml`); the README binds the artifact to `deepseek-ai/DeepSeek-V2` and shows a sample prediction prompt for text ge…

模型作用:DeepSeek-V2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2 is the exact model being wrapped and served. The artifact downloads DeepSeek-V2 weights, initializes vLLM with tensor_parallel_size=8 and max_model_len=8192, and streams generated text from user prompts thro…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an open-source deployment/integration artifact rather than a customer business story. Evidence is strong for exact model binding and a reachable public artifact, but the repository README labels the implementati…

原始记录:lucataco packaged DeepSeek-V2 as a Cog model for Replicate-style deployment

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

gabrielchua / Repo Explainer 使用 Gemini 1.5 Pro (May) 处理软件工程任务执行

gabrielchua / Repo Explainer · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:GitHub 最后复核:2026-06-27T05:47:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

gabrielchua / Repo Explainer 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:47:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:用户输入 GitHub 仓库 URL,应用读取仓库结构、README 和相关代码文件,并用 Gemini-1.5-pro-latest 生成仓库用途、功能和代码结构解释,同时提供交互式问答界面继续追问仓库细节。

公开产物:公开 Streamlit/GitHub 应用 Repo Explainer,README 明确展示了仓库分析、AI 解释和聊天功能,并标注使用 Gemini-1.5-pro-latest;项目还附带 demo.gif 和线上 Streamlit 入口。

模型作用:Gemini 1.5 Pro 的长上下文能力被用于一次性吸收仓库目录、README 和多类代码文件,生成综合摘要并回答针对仓库实现细节的问题。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自开源仓库 README;未核验线上 Streamlit 当前运行状态,artifact 以可访问 GitHub 仓库为准。

原始记录:Repo Explainer 用 Gemini 1.5 Pro 分析并讲解 GitHub 仓库

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

trippedout / Gemini Video Scrubber 使用 Gemini 1.5 Pro (May) 处理多模态内容处理

trippedout / Gemini Video Scrubber · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:GitHub 最后复核:2026-06-27T05:47:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

trippedout / Gemini Video Scrubber 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T05:47:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:应用把用户上传的视频发送到 Gemini File API,并提示 Gemini 返回与用户请求相关的时间戳和描述;前端将这些时间戳变成可点击、可连续播放的视频片段。

公开产物:公开仓库 Gemini Video Scrubber 提供可运行的 Next.js 演示应用、截图和 YouTube walkthrough;README 明确说明它展示 Gemini 1.5 Pro 的视频理解能力,并用于从最长约 1 小时视频中寻找可跨团队分享的片段。

模型作用:Gemini 1.5 Pro 负责理解视频帧和音频内容,抽取满足提示条件的关键时刻,并生成时间戳化描述,应用层负责播放和可视化这些片段。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目 README 称为 demo/prototype,但包含公开可运行代码、截图和视频 walkthrough,且描述了内部用同一技术寻找长视频片段的实际用途。

原始记录:Gemini Video Scrubber 用 Gemini 1.5 Pro 定位长视频片段时间戳

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

di37 / YouTube Notes Generator 使用 Gemini 1.5 Pro (May) 处理多模态内容处理

di37 / YouTube Notes Generator · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:GitHub 最后复核:2026-06-27T05:47:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

di37 / YouTube Notes Generator 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T05:47:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Streamlit 应用接收 YouTube URL,使用音频理解而非单纯字幕,为教程类或课堂讲座类视频生成高质量、全面的笔记;README 明确模型可选 gemini-1.5-pro 或 gemini-1.5-flash。

公开产物:公开仓库包含应用代码、安装运行说明和示例截图,说明用户输入 YouTube URL 后点击 Generate Notes 即可生成详细视频笔记。

模型作用:Gemini 1.5 Pro 被用作核心多模态/音频理解模型,根据视频音频内容和系统提示提炼结构化学习笔记。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:仓库同时支持 Gemini 1.5 Flash;候选绑定到 README 明确列出的 Gemini-1.5-Pro 模型能力路径。

原始记录:YouTube Notes Generator 用 Gemini 1.5 Pro 从视频音频生成学习笔记

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

University of Michigan Herbarium / … 使用 Gemini 1.5 Pro (May) 处理文档理解和结构化处理

University of Michigan Herbarium / Gene-Weaver VoucherVision · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:GitHub 最后复核:2026-06-27T05:47:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

University of Michigan Herbarium / Gene-Weaver VoucherVision 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T05:47:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:VoucherVision 面向自然历史标本标签转录工作流,使用大语言/视觉模型从标本图片中读取标签文字并输出 JSON 或 CSV 等结构化数据;README 更新中明确列出 Google Gemini models(Gemini-1.5-Pro、Gemini-1.5-Flash、Gemini-1.5-Flash-8B)作为 OCR 引擎,并指出 Gemini-1.5-Pro 是首个/唯一能可靠读取 included demo image 中困难草写体的 OCR 引擎。

公开产物:公开仓库提供 VoucherVision Beta 工具、安装和 UI 文档;README 记录 MICH herbarium 已处理约 50,000 张图片,并说明 Gemini-1.5-Pro 在困难草写体 OCR 任务上的实际效果。

模型作用:Gemini 1.5 Pro 作为 OCR/VLM 引擎读取高难度手写和印刷混合的标本标签,为后续字段抽取、CSV/JSON 输出和人工校对提供文本基础。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 后续推荐工作流已更新到 Gemini 2.0 Flash,但历史更新明确记载 Gemini-1.5-Pro 的 OCR 使用和效果;此候选只主张 Gemini 1.5 Pro 在该工作流中的已公开用例。

原始记录:VoucherVision 用 Gemini 1.5 Pro 读取自然历史标本标签并结构化转录

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

arjunprabhulal / DVD Rental Assista… 使用 Gemini 1.5 Pro (May) 处理智能体流程编排

arjunprabhulal / DVD Rental Assistant · Gemini 1.5 Pro (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (May) 来源平台:GitHub 最后复核:2026-06-27T05:47:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

arjunprabhulal / DVD Rental Assistant 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T05:47:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:FastAPI + Streamlit 应用把 Gemini、LlamaIndex AgentWorkflow 和 GenAI Toolbox 结合起来,为 Pagila/DVD rental PostgreSQL 数据库提供自然语言查询、电影搜索、可用性检查、客户记录和租赁状态等操作。README 的实现片段明确初始化 GoogleGenAI(model="gemini-1.5-pro")。

公开产物:公开仓库提供完整应用结构、架构图、后端/前端运行入口和配置说明;用户可在本地启动后通过 Streamlit 前端和 FastAPI 后端对 DVD rental 数据库进行对话式管理。

模型作用:Gemini 1.5 Pro 承担对用户自然语言请求的理解、工具调用意图生成和上下文响应生成;GenAI Toolbox/LlamaIndex 负责把模型决策连接到安全的数据库工具。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该项目是公开工程应用/示例系统,未看到生产客户部署声明;证据满足公开 repo、明确模型、明确任务和可运行产物。

原始记录:DVD Rental Assistant 用 Gemini 1.5 Pro 驱动电影租赁数据库问答与操作代理

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

syaikhipin 使用 Kimi K2 Thinking 处理软件工程任务执行

syaikhipin · Kimi K2 Thinking

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 Thinking 来源平台:github 最后复核:2026-06-27T05:52:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

syaikhipin 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:52:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a public Gradio application that accepts formal or AI-generated text, sends it through the OpenRouter chat-completions API using the moonshotai/Kimi-K2-Thinking model, applies a natural-rewriting prompt, filters t…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository for KIMI-AI-Humanizer with Hugging Face Space metadata, README setup instructions, and app.py code. The README states the app rewrites content using Kimi-K2-Thinking, and app.py sets the default…

模型作用:Kimi K2 Thinking 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2 Thinking is the core generation model for the application: it receives the user's source text and rewriting prompt, produces the naturalized text, and its thinking output is explicitly filtered before the final …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub repository and source code from the maintainer; no usage metrics or customer deployment claims are published. The Hugging Face Space URL advertised by repo metadata returned 401 during collec…

原始记录:syaikhipin built a Gradio human-text rewriting app with Kimi K2 Thinking

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

untanao 使用 Kimi K2 处理软件工程任务执行

untanao · Kimi K2

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 来源平台:github 最后复核:2026-06-27T05:52:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

untanao 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:52:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Open-source web agent where a user pastes an EVM token contract address and the system gathers Etherscan metadata, DexScreener market data, GoPlus security signals, honeypot simulation results, and web evidence to asses…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The application streams tool-call progress to the UI and produces a structured Markdown due-diligence report with a 0-100 risk score and verdict for the submitted token.

模型作用:Kimi K2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states Kimi K2 works out of the box and is the default reasoning model via Fireworks (`accounts/fireworks/models/kimi-k2p6`). Kimi K2 drives the agent loop, selects among available tools, interprets retrieved…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public GitHub README and code artifact; the project is model-agnostic and supports other LLMs, but its documented default model is Kimi K2 via Fireworks.

原始记录:Token DD Agent uses Kimi K2 as the default reasoning model for crypto token due diligence

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yberkayozkan 使用 Kimi K2 处理软件工程任务执行

yberkayozkan · Kimi K2

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 来源平台:github 最后复核:2026-06-27T05:52:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yberkayozkan 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T05:52:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Open-source LangChain/Streamlit assistant for geophysical workflows, including reading and analyzing SEGY seismic data, processing LAS well logs, answering natural-language questions, and generating seismic sections, we…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents an interactive chat interface that can answer seismic and well-data questions such as showing inline/crossline sections, generating amplitude histograms, listing wells, and plotting logs.

模型作用:Kimi K2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states the assistant is powered by the Kimi-K2 model through OpenRouter (`moonshotai/kimi-k2:free`). Kimi-K2 provides natural-language understanding, intelligent tool selection and parameter extraction, and c…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public GitHub README and code artifact; runtime requires an OpenRouter API key and user-provided geophysical files, but the model binding is explicitly documented as Kimi-K2.

原始记录:Seismo-Lingo uses Kimi-K2 for natural-language geophysical data analysis

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LLM-Red-Team 使用 Kimi K2 处理软件工程任务执行

LLM-Red-Team · Kimi K2

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 来源平台:github 最后复核:2026-06-27T06:01:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LLM-Red-Team 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a public command-line integration that lets developers drive Claude Code through Moonshot's Kimi service, explicitly using the kimi-k2-0711-preview model and the Kimi Open Platform API key flow.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository for Kimi CC with install script, multilingual README files, documented Kimi Open Platform API-key setup, and a workflow where running the claude command is backed by Kimi K2 0711 Preview for low…

模型作用:Kimi K2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2 0711 Preview is the model backend that supplies the coding-agent reasoning and generation used through Claude Code; the project wraps authentication and endpoint configuration so the Claude Code interface can de…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub README and repository explicitly naming kimi-k2-0711-preview. It is a community developer tool rather than an official Moonshot customer story; no independent usage metrics are published.

原始记录:LLM-Red-Team built Kimi CC to run Claude Code on Kimi K2 0711 Preview

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

PratiekSonare 使用 DeepSeek V3 0324 处理软件工程任务执行

PratiekSonare · DeepSeek V3 0324

A
厂商:DeepSeek 模型:DeepSeek V3 0324 来源平台:github 最后复核:2026-06-27T06:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

PratiekSonare 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Connect a user's daily tasks, durations, and boredom scores to an LLM so the application can produce an optimized timetable with Pomodoro-style break recommendations.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public C++/GTK repository; NoteEditor.cpp posts to OpenRouter with model deepseek/deepseek-chat-v3-0324:free and requests JSON schedule items containing time_slot, task, duration, technique_applied, and notes.

模型作用:DeepSeek V3 0324 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 0324 transforms structured task inputs into an optimized daily schedule and explains the work/break technique for each task.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:No README was available at the default branch path; evidence is the reachable source file plus the public repository description.

原始记录:Desktop daily schedule optimizer using DeepSeek V3 0324 for Pomodoro planning

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

5ire / nanbingxyz 使用 Claude Instant 处理智能体流程编排

5ire / nanbingxyz · Claude Instant

A
厂商:Anthropic / Claude 模型:Claude Instant 来源平台:GitHub 最后复核:2026-06-27T06:02:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

5ire / nanbingxyz 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:02:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:5ire is a cross-platform desktop AI assistant and MCP client; its model handling code explicitly recognizes claude-instant-1 as the Claude 1/Instant option users can select for assistant conversations and tool/MCP workf…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public desktop AI assistant artifact that integrates Anthropic Claude Instant as one of its supported chat models for end-user assistant interactions.

模型作用:Claude Instant 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Claude Instant supplies the conversational language model behind 5ire sessions when the claude-instant-1 option is selected, enabling low-latency assistant responses and tool-driven workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source code evidence verifies model binding; no private production telemetry is claimed.

原始记录:5ire exposes Claude Instant in a desktop AI assistant and MCP client

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Ahnuf-Karim-Chowdhury 使用 DeepSeek V3 0324 处理软件工程任务执行

Ahnuf-Karim-Chowdhury · DeepSeek V3 0324

A
厂商:DeepSeek 模型:DeepSeek V3 0324 来源平台:github 最后复核:2026-06-27T06:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Ahnuf-Karim-Chowdhury 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Streamlit app where users upload PDF or TXT resumes and optionally a target role, then receive feedback on structure, clarity, content impact, and job fit.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository and README describing an app that extracts resume text, sends it to DeepSeek V3 0324 through OpenRouter, and returns balanced, specific, actionable resume feedback.

模型作用:DeepSeek V3 0324 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 0324 performs the resume critique and recommendation generation after the app parses the uploaded document into prompt text.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The linked Streamlit deployment redirected to an auth flow during verification, so the repository is used as the accessible artifact URL.

原始记录:Panda Resume Critiquer uses DeepSeek V3 0324 to analyze resumes and return actionable feedback

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Julep AI 使用 Claude Instant 处理智能体流程编排

Julep AI · Claude Instant

A
厂商:Anthropic / Claude 模型:Claude Instant 来源平台:GitHub 最后复核:2026-06-27T06:02:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Julep AI 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:02:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Julep is a public agent platform; its agents API registry lists claude-instant-1 and claude-instant-1.2 with 100k-token context windows in the Claude model set used by the platform.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public AI-agent platform artifact where developers can build agent workflows against registered Claude Instant model identifiers.

模型作用:Claude Instant 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Claude Instant contributes the LLM backend option for Julep agents, enabling agent conversations and task execution when that registered model is chosen.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the platform model registry rather than a named customer deployment; the repo itself is the public artifact.

原始记录:Julep includes Claude Instant in its AI agent platform model registry

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

tiankongzhise 使用 DeepSeek V3 0324 处理软件工程任务执行

tiankongzhise · DeepSeek V3 0324

A
厂商:DeepSeek 模型:DeepSeek V3 0324 来源平台:github 最后复核:2026-06-27T06:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tiankongzhise 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Classify keyword lists to determine whether each keyword is an artificial SEO-created term, reducing the difficulty of keyword screening.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public Python repository with README stating that deepseek_v3 0324 judges whether keywords are SEO-created words; source code reads keyword.txt, calls an OpenAI-compatible endpoint for AI responses, and stores keyword S…

模型作用:DeepSeek V3 0324 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 0324 supplies the semantic judgement for each keyword, producing classification scores that the application persists for later filtering.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The README explicitly names deepseek_v3 0324, but the checked source uses a provider endpoint ID rather than a self-describing model slug.

原始记录:SEO keyword classifier using DeepSeek V3 0324 to identify artificially generated SEO terms

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

RA.Aid / ai-christianson 使用 Claude Instant 处理软件工程任务执行

RA.Aid / ai-christianson · Claude Instant

A
厂商:Anthropic / Claude 模型:Claude Instant 来源平台:GitHub 最后复核:2026-06-27T06:02:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

RA.Aid / ai-christianson 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:02:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RA.Aid is an autonomous software-development agent; its callback/cost handling table explicitly includes claude-instant-1, supporting runs where Claude Instant is selected for coding-agent work.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public coding-agent artifact that can account for and run Anthropic Claude Instant as an LLM choice during autonomous research, planning, and implementation tasks.

模型作用:Claude Instant 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Instant provides a lower-cost, faster Claude-family model option for RA.Aid agent steps, contributing natural-language reasoning and code-task assistance when configured.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence verifies model support in a real agent tool; it does not claim Claude Instant is the current default model.

原始记录:RA.Aid tracks Claude Instant for autonomous software-development agent usage

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AgentPilot / jbexta 使用 Claude Instant 处理智能体流程编排

AgentPilot / jbexta · Claude Instant

A
厂商:Anthropic / Claude 模型:Claude Instant 来源平台:GitHub 最后复核:2026-06-27T06:02:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AgentPilot / jbexta 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:02:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AgentPilot is a workflow-automation platform for AI-driven tasks; its default/reset model catalog includes claude-instant-1 and claude-instant-1.2 as Anthropic chat model choices.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public workflow automation application where users can configure Claude Instant as the chat model for single-LLM chats or multi-member AI workflows.

模型作用:Claude Instant 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Claude Instant contributes the model execution layer for AgentPilot chat and workflow nodes when selected, powering assistant responses within automated workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The evidence is the shipped model catalog in an open-source app; no usage-volume metric is asserted.

原始记录:AgentPilot ships Claude Instant options for AI workflow automation

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

eunomia-bpf 使用 Claude Instant 处理智能体流程编排

eunomia-bpf · Claude Instant

A
厂商:Anthropic / Claude 模型:Claude Instant 来源平台:GitHub 最后复核:2026-06-27T06:02:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

eunomia-bpf 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:02:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:GPTtrace is an experiment/tool for generating eBPF programs and Linux tracing workflows from natural language; its execution code calls LiteLLM completion with model="claude-instant-1".

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public developer tool that turns natural-language tracing requests into eBPF/tracing outputs, with Claude Instant explicitly configured as the completion model in the execution path.

模型作用:Claude Instant 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Claude Instant performs the natural-language-to-code generation step, producing responses that GPTtrace uses to create or explain eBPF tracing programs.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository labels the project experimental and not for production use, but it is a concrete public artifact rather than a tutorial or benchmark.

原始记录:GPTtrace uses Claude Instant to generate eBPF tracing programs from natural language

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DayuanJiang / Next AI Draw.io 使用 DeepSeek R1 0528 处理真实任务执行

DayuanJiang / Next AI Draw.io · DeepSeek R1 0528

A
厂商:DeepSeek 模型:DeepSeek R1 0528 来源平台:github 最后复核:2026-06-27T05:58:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DayuanJiang / Next AI Draw.io 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T05:58:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Next AI Draw.io 是一个 AI-powered diagram creation tool,用自然语言创建、修改和增强 draw.io 图表;项目的 EdgeOne chat completions 兼容接口显式允许 deepseek-r1-0528 别名。

公开产物:公开开源的 Next.js 图表生成应用和在线 demo;README 描述其产物为通过自然语言命令生成和编辑 diagrams。

模型作用:edge-functions/api/edgeai/chat/completions.ts 将 deepseek-r1-0528 映射到 @tx/deepseek-ai/deepseek-r1-0528,并将其列入 ALLOWED_MODELS,使该模型可作为图表生成/编辑对话的推理后端。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自应用 README 与模型路由源码;不是模型集合页,模型被接入具体图表生成产品路径。

原始记录:Next AI Draw.io 将 DeepSeek R1 0528 接入自然语言图表生成工具

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Hrishikesh332 / Video2Game 使用 DeepSeek R1 0528 处理可玩交互原型构建

Hrishikesh332 / Video2Game · DeepSeek R1 0528

A
厂商:DeepSeek 模型:DeepSeek R1 0528 来源平台:github 最后复核:2026-06-27T05:58:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Hrishikesh332 / Video2Game 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T05:58:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Video2Game 接收 YouTube 教学/讲解视频,先用 TwelveLabs 做视频理解,再用 SambaNova 上的 DeepSeek 生成可交互学习游戏代码。

公开产物:公开 GitHub 项目提供从教育视频到 interactive games 的应用代码、架构图和 generated_games 输出目录设计。

模型作用:backend/config.py 将 SAMBANOVA_MODEL 固定为 DeepSeek-R1-0528;README 说明游戏生成环节由 SambaNova - Deepseek 完成,因此该模型负责把视频分析结果转化为游戏逻辑和代码。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README 与后端配置;这是应用仓库中的实际生成流程,不是 leaderboard 或单纯教程。

原始记录:Video2Game 使用 DeepSeek R1 0528 将教学视频生成互动游戏

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

PandaAI-Tech / PandaAI QuantFlow 使用 DeepSeek R1 0528 处理研究分析和报告生成

PandaAI-Tech / PandaAI QuantFlow · DeepSeek R1 0528

A
厂商:DeepSeek 模型:DeepSeek R1 0528 来源平台:github 最后复核:2026-06-27T05:58:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

PandaAI-Tech / PandaAI QuantFlow 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T05:58:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:PandaAI QuantFlow 是面向量化研究、策略开发、因子分析和回测的可视化工作流平台,内置 LLM 服务配置以支持平台中的智能辅助能力。

公开产物:公开开源的量化交易和机器学习工作流平台,README 展示了可视化工作流编排、策略回测、金融数据分析和机器学习应用等产物形态。

模型作用:llm_model_type.py 中 DeepSeek-R1 使用 DeepSeek API 的 deepseek-reasoner,并将 version 标注为 DeepSeek-R1-0528;模型为量化工作流平台提供推理型 LLM 后端。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自平台 README 与 LLM 模型配置;DeepSeek API 的 deepseek-reasoner 版本绑定为 R1-0528。

原始记录:PandaAI QuantFlow 将 DeepSeek R1 0528 用于量化工作流平台的 LLM 能力

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kevgeoleo 使用 DeepSeek V3 0324 处理软件工程任务执行

kevgeoleo · DeepSeek V3 0324

A
厂商:DeepSeek 模型:DeepSeek V3 0324 来源平台:github 最后复核:2026-06-27T06:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kevgeoleo 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a retrieval-augmented chatbot that embeds a local dataset, retrieves relevant chunks, and answers user questions such as PAN-card FAQ or instruction-manual queries.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository with chatbot.py using model deepseek/deepseek-chat-v3-0324:free and README instructions for running the RAG chatbot over a dataset.

模型作用:DeepSeek V3 0324 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 0324 is the generation model called through OpenRouter; retrieved context is inserted into the prompt and the model produces context-aware answers.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small personal repository; evidence is the project README and source code, not a third-party customer story.

原始记录:RAG chatbot answering dataset-specific questions with DeepSeek V3 0324 via OpenRouter

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

FaustoS88 使用 DeepSeek V3 0324 处理软件工程任务执行

FaustoS88 · DeepSeek V3 0324

A
厂商:DeepSeek 模型:DeepSeek V3 0324 来源平台:github 最后复核:2026-06-27T06:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

FaustoS88 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a conversational AI companion that transcribes speech with Faster-Whisper, stores long-term memories with Mem0 and PostgreSQL/pgvector, and answers with remembered context.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository containing STT-mem0 and STT-mem0V2 implementations; README documents DeepSeek/OpenRouter support and the code sets llm_model to deepseek/deepseek-chat-v3-0324:free for OpenRouter.

模型作用:DeepSeek V3 0324 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek V3 0324 is the chat LLM that produces assistant replies after the application retrieves recent conversation and stored memory context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Personal open-source project; README also supports the DeepSeek API alias deepseek-chat, while the OpenRouter path explicitly pins deepseek-chat-v3-0324:free.

原始记录:Speech-to-text AI companion with long-term memory using DeepSeek V3 0324

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

com-wuqi 使用 DeepSeek V3.1 Terminus 处理软件工程任务执行

com-wuqi · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T09:38:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

com-wuqi 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Generate and publish a FastAPI, SQLite, SQLModel, JWT, and bcrypt based user registration, login, current-user, and admin user-management backend service.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains the completed FastAPI authentication server with app modules, route files, startup script, API test script, dependency list, README, and documented endpoints including /auth/register, /aut…

模型作用:DeepSeek V3.1 Terminus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The repository description explicitly says the project was completed by Claude Code with SiliconFlow and DeepSeek-V3.1-Terminus; the model is therefore tied to code-generation and implementation assistance for the backe…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a GitHub repository description plus the published code artifact; model usage is self-reported by the repository owner and includes Claude Code/SiliconFlow in the toolchain, so contribution is assistant-assi…

原始记录:com-wuqi published a FastAPI user-management and authentication service built with DeepSeek V3.1 Terminus through SiliconFlow

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Zhichii / RikkaHub 使用 DeepSeek V3.1 Terminus 处理真实任务执行

Zhichii / RikkaHub · DeepSeek V3.1 Terminus

A
厂商:DeepSeek 模型:DeepSeek V3.1 Terminus 来源平台:github 最后复核:2026-06-27T06:03:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Zhichii / RikkaHub 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T06:03:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:RikkaHub 用户在 issue 中列出 deepseek-ai/DeepSeek-V3.1-Terminus 和 Pro/deepseek-ai/DeepSeek-V3.1-Terminus 支持 enable_thinking,并在 Android 版 RikkaHub 1.6.16 中测试 Thinking 按钮/API 格式兼容性。

公开产物:Issue 记录了可复现的问题和 workaround:手动添加 Claude/Anthropic 格式的 SiliconFlow 或 DeepSeek API 后 thinking 参数可用;页面还附有展示硅基流动和 DeepSeek 方案的 mp4 截图证据。

模型作用:DeepSeek-V3.1-Terminus 是被集成到移动端聊天客户端的 thinking-capable 模型,用于产生带 thinking/reasoning 行为的对话输出;案例验证了其 API 参数在客户端中的实际适配路径。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自 GitHub issue 和附件,属于真实用户集成/调试场景;不是正式 customer story,但包含具体用户、模型、任务、结果和公开 artifact。

原始记录:RikkaHub 用户验证 DeepSeek-V3.1-Terminus 在 SiliconFlow/DeepSeek API 下的 thinking 参数支持

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

sansan0 / TrendRadar 使用 Gemini 1.5 Pro (Sep) 处理软件工程任务执行

sansan0 / TrendRadar · Gemini 1.5 Pro (Sep)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (Sep) 来源平台:github 最后复核:2026-06-27T06:09:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

sansan0 / TrendRadar 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:09:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:TrendRadar is an open-source public-opinion and trend-monitoring application that aggregates multi-platform hot topics and RSS feeds, then uses an AI provider configured via AI_MODEL; its README lists Google Gemini with…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact is a working self-hosted monitoring system that can generate AI-screened news, Chinese translations for foreign RSS sources, AI analysis briefs, and push notifications to WeChat, Feishu, DingTalk, Te…

模型作用:Gemini 1.5 Pro (Sep) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Pro supplies the large-language-model layer for classifying and screening information, translating feed content, and writing trend-analysis summaries before alerts are sent.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the project README and binds to the Gemini 1.5 Pro API/model name; it is an open-source integration rather than a Google customer story.

原始记录:TrendRadar uses Gemini 1.5 Pro for AI public-opinion monitoring, translation, and alert summaries

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Sakana AI 使用 Gemini 1.5 Pro (Sep) 处理软件工程任务执行

Sakana AI · Gemini 1.5 Pro (Sep)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (Sep) 来源平台:github 最后复核:2026-06-27T06:09:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Sakana AI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:09:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The AI-Scientist repository is a public system for automated open-ended scientific discovery; its README documents Google Gemini support through the google-generativeai library and explicitly names gemini-1.5-pro as a s…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact provides code for generating research ideas, running experiments, writing papers, and using supporting services such as Semantic Scholar, with Gemini 1.5 Pro available as one of the model backends for those…

模型作用:Gemini 1.5 Pro (Sep) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Pro acts as the LLM backend that can drive scientific ideation, experiment planning, code/text generation, and paper-writing steps in the AI-Scientist workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public research-system repository and README model configuration; individual generated papers may depend on the operator's chosen model at run time.

原始记录:Sakana AI AI-Scientist supports Gemini 1.5 Pro for automated scientific discovery workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

duty1g / SubCat 使用 Gemini 1.5 Pro (Sep) 处理研究分析和报告生成

duty1g / SubCat · Gemini 1.5 Pro (Sep)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (Sep) 来源平台:github 最后复核:2026-06-27T06:09:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

duty1g / SubCat 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T06:09:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SubCat is a public passive subdomain-discovery tool for security professionals and bug-bounty hunters. Its README documents an --ai mode and shows a concrete command using --ai-model gemini-1.5-pro to analyze 403 Forbid…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact is a CLI tool that discovers subdomains from many passive sources and can produce AI-assisted explanations or diagnostics for blocked/error responses during reconnaissance.

模型作用:Gemini 1.5 Pro (Sep) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Pro provides the natural-language analysis layer that interprets HTTP error/status-code situations and helps users understand failures in the security-recon workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub README with an explicit gemini-1.5-pro command example; runtime use requires the operator to supply a Gemini API key.

原始记录:SubCat uses Gemini 1.5 Pro for AI analysis of passive subdomain-discovery errors

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

coffeegrind123 / gemini-for-claude-… 使用 Gemini 1.5 Pro (Sep) 处理软件工程任务执行

coffeegrind123 / gemini-for-claude-code · Gemini 1.5 Pro (Sep)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (Sep) 来源平台:github 最后复核:2026-06-27T06:09:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

coffeegrind123 / gemini-for-claude-code 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:09:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:gemini-for-claude-code is a Python proxy that lets Claude Code-compatible clients call Google's Gemini models. Its README configuration maps BIG_MODEL to gemini-1.5-pro-latest for sonnet/opus-style requests and SMALL_MO…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact is a runnable proxy server that accepts Claude Code-style traffic, performs model mapping, and streams or returns Gemini responses through the configured Gemini API key.

模型作用:Gemini 1.5 Pro (Sep) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Pro Latest is the high-capability backend model used to answer the larger coding-assistant requests that the proxy receives from Claude Code-compatible tooling.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public integration repository; gemini-1.5-pro-latest is an alias in the Gemini 1.5 Pro family and may resolve according to Google's current API routing.

原始记录:gemini-for-claude-code maps Claude Code big-model requests to Gemini 1.5 Pro Latest

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

milanglacier / minuet-ai.nvim 使用 Gemini 1.5 Pro (Sep) 处理文档理解和结构化处理

milanglacier / minuet-ai.nvim · Gemini 1.5 Pro (Sep)

A
厂商:Google / Gemini 模型:Gemini 1.5 Pro (Sep) 来源平台:github 最后复核:2026-06-27T06:09:07Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

milanglacier / minuet-ai.nvim 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T06:09:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Minuet AI is a Neovim plugin for AI code completion. Its README documents the Minuet change_model command and gives a concrete example, Minuet change_model gemini:gemini-1.5-pro-latest, for switching the active completi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact is an installable editor plugin that provides as-you-type AI code completions and lets users switch providers/models interactively or by command, including Gemini 1.5 Pro Latest.

模型作用:Gemini 1.5 Pro (Sep) 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.5 Pro Latest serves as the code-generation and completion backend that predicts code continuations inside the Neovim editor workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the project README and a concrete model-switching example; actual completions depend on user configuration and API credentials.

原始记录:Minuet AI Neovim plugin uses Gemini 1.5 Pro Latest as a selectable code-completion model

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

E2B (e2b-dev) 使用 GPT-4 Turbo 处理软件工程任务执行

E2B (e2b-dev) · GPT-4 Turbo

A
厂商:OpenAI 模型:GPT-4 Turbo 来源平台:GitHub 最后复核:2026-06-27T06:24:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

E2B (e2b-dev) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AI Developer is an agent for working on a user's GitHub repository: reading files, writing code, making pull requests, pulling repositories, responding to feedback, and running commands in an E2B sandbox.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides an installable AI developer artifact that performs repository changes and command execution in isolated E2B cloud sandboxes.

模型作用:GPT-4 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly titles the project as powered by GPT-4-Turbo and OpenAI's Assistants API; GPT-4 Turbo supplies the agent reasoning and natural-language task execution while E2B supplies the sandbox runtime.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository README; no independent production customer metrics are provided.

原始记录:E2B AI Developer uses GPT-4 Turbo and the Assistants API to modify GitHub repositories inside sandboxes

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Aymen Furter 使用 GPT-4 Turbo 处理软件工程任务执行

Aymen Furter · GPT-4 Turbo

A
厂商:OpenAI 模型:GPT-4 Turbo 来源平台:GitHub 最后复核:2026-06-27T06:24:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Aymen Furter 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Microagents dynamically creates small, task-specific agents in response to user-assigned tasks, assesses their functional performance, and iteratively improves the agents.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes a working Python framework with demo instructions for creating self-improving microservice-sized agents.

模型作用:GPT-4 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README's Built With and prerequisites sections name OpenAI's GPT-4 Turbo as the LLM used for deducing task execution methods and driving agent generation alongside text-embedding-ada-002.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source project evidence; README binds the model, but it does not report external deployment scale.

原始记录:Microagents uses GPT-4 Turbo to generate and assess self-improving task-specific agents

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Bhautik (yesbhautik) 使用 GPT-4 Turbo 处理软件工程任务执行

Bhautik (yesbhautik) · GPT-4 Turbo

A
厂商:OpenAI 模型:GPT-4 Turbo 来源平台:GitHub 最后复核:2026-06-27T06:24:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Bhautik (yesbhautik) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Telegram bot for general assistant conversations, group chat support, code assistant mode, voice message recognition, DALL-E image generation mode, and multiple specialized chat modes.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides the bot implementation and README demonstration of a Telegram-based AI assistant with fast streaming responses and multiple modes.

模型作用:GPT-4 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states the bot is powered by GPT-4 Turbo; GPT-4 Turbo provides the assistant responses, code help, role-based chat modes, and conversation intelligence behind the Telegram interface.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is README-based and does not include usage telemetry; still a concrete public bot artifact with explicit GPT-4 Turbo binding.

原始记录:Master AI BOT uses GPT-4 Turbo to power a Telegram assistant with streaming, roles, and voice features

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

small-cactus 使用 GPT-4 Turbo 处理软件工程任务执行

small-cactus · GPT-4 Turbo

A
厂商:OpenAI 模型:GPT-4 Turbo 来源平台:GitHub 最后复核:2026-06-27T06:24:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

small-cactus 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Command-line interface that lets users operate a terminal through conversational natural-language requests, including command interpretation and executing up to three commands simultaneously.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes MagicTerminal, a CLI application designed to make terminal operations more intuitive without memorizing command syntax.

模型作用:GPT-4 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README says MagicTerminal integrates OpenAI's GPT-4-Turbo model for natural language processing; the model interprets user intent and converts it into command-line instructions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public repo and README provide exact model binding; production adoption metrics are not provided.

原始记录:MagicTerminal uses GPT-4 Turbo to translate natural language into terminal commands

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Arjun (4arjun) 使用 GPT-4 Turbo 处理软件工程任务执行

Arjun (4arjun) · GPT-4 Turbo

A
厂商:OpenAI 模型:GPT-4 Turbo 来源平台:GitHub 最后复核:2026-06-27T06:24:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Arjun (4arjun) 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Backend service for an AI voice/chat companion that helps users practice and improve language skills through interactive conversations with session history and supportive feedback.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public backend repository includes the GPT-4 Turbo chat flow and documentation for a language-practice companion application.

模型作用:GPT-4 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README and sample code explicitly call OpenAI GPT-4 Turbo / model="gpt-4-turbo"; GPT-4 Turbo generates the companion's conversational responses and feedback while maintaining context through conversation history.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public backend repository; the deployed frontend or user metrics are not separately verified.

原始记录:ConvoVoice backend uses GPT-4 Turbo for a friendly language-practice companion

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Vitality Squad 使用 PALM-2 处理软件工程任务执行

Vitality Squad · PALM-2

A
厂商:Google / Gemini 模型:PALM-2 来源平台:github 最后复核:2026-06-27T06:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Vitality Squad 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DermDetect is a Streamlit skin-disease app where patients can discuss skin-related medical concerns with a chatbot and submit images for disease classification across seven categories.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository README reports an image-classification model trained on about 15,000 Kaggle images and approximately 86% classification accuracy, paired with an interactive PaLM 2 chatbot interface for user guidance.

模型作用:PALM-2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:PaLM 2 powers the conversational chatbot portion of DermDetect, giving users an interface for discussing concerns and understanding results alongside the separate image classifier.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public GitHub project with README and screenshot; medical-use claims are project-reported and should not be treated as clinically validated. Evidence directly names PaLM 2.

原始记录:Vitality Squad built DermDetect with a PaLM 2 chatbot for skin-health conversations

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Poorna-Chandra-D / QueryGenie proje… 使用 PALM-2 处理软件工程任务执行

Poorna-Chandra-D / QueryGenie project · PALM-2

A
厂商:Google / Gemini 模型:PALM-2 来源平台:github 最后复核:2026-06-27T06:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Poorna-Chandra-D / QueryGenie project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:QueryGenie lets business users ask database questions in natural language, such as inventory value or discounted revenue questions, and receive answers without writing SQL.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository describes a Streamlit application that translates plain-English questions into SQL via LangChain SQLDatabaseChain and few-shot prompting, then returns precise answers for retail/logistics/e-commerce style…

模型作用:PALM-2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Google PaLM 2 is the LLM used through LangChain to interpret natural-language questions, generate SQL, and orchestrate the query-answering workflow with Hugging Face embeddings and ChromaDB for semantic examples.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public repository with explicit PaLM 2 mention and concrete app task; appears to be an individual developer project/demo rather than a deployed enterprise customer story.

原始记录:QueryGenie used Google PaLM 2 to turn plain-English business questions into SQL answers

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

keanteng / Streamlit Chatbot project 使用 PALM-2 处理软件工程任务执行

keanteng / Streamlit Chatbot project · PALM-2

A
厂商:Google / Gemini 模型:PALM-2 来源平台:github 最后复核:2026-06-27T06:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

keanteng / Streamlit Chatbot project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project builds a Streamlit chatbot that reads input data and recommends Netflix movies or shows, with the README noting that the same pattern can apply to jobs, movies, songs, or related text data.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides a runnable Streamlit application and README instructions for launching a PaLM-2-powered recommendation chatbot locally.

模型作用:PALM-2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The PaLM-2 API supplies the large-language-model reasoning and conversational response layer that turns user prompts and input data into recommendations.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public GitHub repo explicitly names the PaLM-2 API and a concrete recommendation task; small individual/demo project and not a production customer deployment.

原始记录:A Streamlit chatbot used the PaLM-2 API to recommend Netflix movies and shows from input data

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Sakana AI 使用 o1-preview 处理研究分析和报告生成

Sakana AI · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:GitHub 最后复核:2026-06-27T09:41:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Sakana AI 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T09:41:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:AI Scientist-v2 是 Sakana AI 开源的自动科研系统,README 的示例运行命令明确将 --model_writeup 设置为 o1-preview-2024-09-12,用于在实验阶段后生成研究论文 writeup。

公开产物:产物是公开可运行的 AI Scientist-v2 仓库;运行流程会在 experiments 目录生成实验日志、tree visualization,并在 writeup 阶段产出 timestamp_ideaname.pdf 论文文件。

模型作用:o1-preview 绑定在 writeup 阶段,负责把自动实验结果组织成科研论文草稿/报告,是最终公开科研产物生成链路中的核心写作推理模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README 的示例命令,明确到 dated model o1-preview-2024-09-12;未声称所有默认运行都使用该模型。

原始记录:Sakana AI 在 AI Scientist-v2 中使用 o1-preview 生成科研论文 writeup

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Integuru-AI 使用 o1-preview 处理软件工程任务执行

Integuru-AI · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:GitHub 最后复核:2026-06-27T09:41:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Integuru-AI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:41:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Integuru 是公开开源的集成生成 Agent:用户录制 HAR 后输入目标动作(例如下载 utility bills),系统分析并逆向平台内部 API 请求,然后生成可执行 Python 代码完成该动作。README 明确说明如果用户 OpenAI 账户可用,Integuru 会自动切换到 o1-preview 做 code generation。

公开产物:公开产物是 Integuru GitHub 仓库及其命令行工具;输出结果是命中目标平台内部端点的 runnable Python code,用来执行用户描述的集成动作。

模型作用:o1-preview 被用于代码生成阶段,把已分析出的请求图和用户目标动作转化为可运行的 API 调用代码,承担复杂推理与代码合成角色。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自 README;o1-preview 的使用条件是用户账户中可用,图生成阶段推荐 gpt-4o,代码生成阶段才自动切换到 o1-preview。

原始记录:Integuru 自动切换到 o1-preview 生成调用内部 API 的集成代码

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Agent Laboratory project / Samuel S… 使用 o1-preview 处理软件工程任务执行

Agent Laboratory project / Samuel Schmidgall · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:github 最后复核:2026-06-27T06:28:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Agent Laboratory project / Samuel Schmidgall 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:28:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Agent Laboratory is an end-to-end autonomous research assistant that collects and analyzes papers, plans experiments, prepares data, runs code, and generates a comprehensive report.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact is a public research-agent codebase that orchestrates arXiv, Hugging Face, Python, and LaTeX tools to produce experiment outputs and reports.

模型作用:o1-preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README lists o1-preview as a supported OpenAI backend selectable with the `--llm-backend` flag, so o1-preview can drive the agents' literature analysis, planning, experimentation, and report-generation steps.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence shows concrete product integration and selectable model support; it does not provide a named third-party customer deployment.

原始记录:Agent Laboratory offers o1-preview-backed autonomous research workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

rag-web-ui 使用 MiniMax-M2 处理软件工程任务执行

rag-web-ui · MiniMax-M2

A
厂商:MiniMax 模型:MiniMax-M2 来源平台:GitHub 最后复核:2026-06-27T06:32:37.255041Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

rag-web-ui 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:32:37.255041Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RAG Web UI is an open-source retrieval-augmented generation web application for document management, chunking, vector retrieval, multi-turn chat, and citation-backed knowledge-base Q&A. Its MiniMax provider configuratio…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public web UI/repository that can ingest PDFs, DOCX, Markdown, and text files, build a vector knowledge base, and answer user questions through a chat interface with retrieved references.

模型作用:MiniMax-M2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2 provides the generation layer for MiniMax deployments: it consumes retrieved document context and conversation history to produce grounded answers in the RAG chat workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public implementation/configuration and README, not an external customer story; artifact URL was reachable during collection.

原始记录:RAG Web UI uses MiniMax-M2.7 for knowledge-base question answering

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Y-Research-SBU 使用 MiniMax-M2 处理软件工程任务执行

Y-Research-SBU · MiniMax-M2

A
厂商:MiniMax 模型:MiniMax-M2 来源平台:GitHub 最后复核:2026-06-27T06:32:37.255041Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Y-Research-SBU 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:32:37.255041Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:QuantAgent is a Flask/LangGraph multi-agent system for technical trading analysis. Its web interface defines MiniMax and MiniMax CN providers, validates MiniMax with MiniMax-M2.7-highspeed, and sets MiniMax-M2.7 as the …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public application that lets users select assets, timeframes, and date ranges, then generates technical indicators, chart-pattern analysis, trend reports, and a final trade decision through its web interface and demo …

模型作用:MiniMax-M2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2 acts as the LLM backend for the trading agents, interpreting market/chart context and producing analysis reports and final decision text when the MiniMax provider is selected.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the project code and README for an open-source app; no live hosted deployment was required for eligibility; artifact URL was reachable during collection.

原始记录:QuantAgent uses MiniMax-M2.7 for multi-agent trading analysis

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mrle0429 使用 MiniMax-M2 处理软件工程任务执行

mrle0429 · MiniMax-M2

A
厂商:MiniMax 模型:MiniMax-M2 来源平台:GitHub / web demo 最后复核:2026-06-27T06:32:37.255041Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mrle0429 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:32:37.255041Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:PACT is an open-source pipeline and interactive web demo for constructing mixed human/AI text samples: input text is split into sentences, selected sentences are rewritten by a chosen model, rewritten sentences are back…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A reachable public demo at pact.mrlepro.com plus a repository pipeline that outputs mixed_text and sentence/document labels for AI-text detection or provenance experiments.

模型作用:MiniMax-M2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2 performs the sentence-level rewrite step, creating AI-authored replacement sentences that are inserted back into the source document before label computation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public demo and repository were reachable; evidence comes from README deployment/usage instructions and supported-model configuration rather than a third-party testimonial.

原始记录:PACT exposes MiniMax-M2.7 in a public text-rewriting dataset-construction demo

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yoshiko-pg 使用 o3 处理软件工程任务执行

yoshiko-pg · o3

A
厂商:OpenAI 模型:o3 来源平台:GitHub 最后复核:2026-06-27T06:28:04Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yoshiko-pg 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:28:04Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Public MCP server that lets Claude Code or other AI coding agents consult OpenAI o3 with web search for debugging, latest-library lookup, design review, and complex software-engineering tasks.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides an installable MCP server and o3-search tool; its code defaults OPENAI_MODEL to o3 and calls the OpenAI Responses API with web_search_preview to return researched answers to the host agent.

模型作用:o3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o3 supplies the high-reasoning, web-search-backed consultation layer used by the MCP tool to investigate errors, current documentation, and design questions that the primary agent cannot resolve alone.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository supports overriding the model via environment variable, but the published default and README use case are explicitly bound to OpenAI o3.

原始记录:o3-search-mcp uses OpenAI o3 as an MCP web-search and reasoning assistant for coding agents

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MemTensor 使用 MiniMax-M2 处理软件工程任务执行

MemTensor · MiniMax-M2

A
厂商:MiniMax 模型:MiniMax-M2 来源平台:GitHub 最后复核:2026-06-27T06:32:37.255041Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MemTensor 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:32:37.255041Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MemOS is an open-source memory operating system for AI applications with a self-hosted REST API. Its MiniMax configuration sets MOS_CHAT_MODEL to MiniMax-M2.7 by default and wires MiniMax API key/base URL settings into …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public self-hosted service/repository that can add user messages, store and retrieve memories, and run memory-centric agent workflows through CLI/REST endpoints while selecting MiniMax as the chat backend.

模型作用:MiniMax-M2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2 supplies the chat/reasoning model used by MemOS when the MiniMax provider is configured, enabling memory extraction, response generation, and agent interactions over persisted user context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from public configuration and README provider support; artifact URL was reachable during collection.

原始记录:MemOS supports MiniMax-M2.7 as the chat model for self-hosted memory services

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

grandcanyonsmith 使用 o3 处理软件工程任务执行

grandcanyonsmith · o3

A
厂商:OpenAI 模型:o3 来源平台:GitHub 最后复核:2026-06-27T06:28:04Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

grandcanyonsmith 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:28:04Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:FastAPI service for streaming deep research reports from a /research endpoint, deployable to Render with an OPENAI_API_KEY and a topic query parameter.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains a Docker/Render-ready web service; server.py calls client.responses.stream with model="o3-deep-research" and web_search_preview to emit structured, cited research text.

模型作用:o3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI o3-deep-research performs the long-form reasoning, web search, citation gathering, and synthesis that turns a user-supplied topic into a research report.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This case is bound to the o3-deep-research variant rather than the base chat model; it remains within the OpenAI o3 model family and is explicitly named in code and README.

原始记录:deep-researcher-render exposes an OpenAI o3-deep-research web service on Render

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Hmbown 使用 MiniMax-M2 处理智能体流程编排

Hmbown · MiniMax-M2

A
厂商:MiniMax 模型:MiniMax-M2 来源平台:GitHub 最后复核:2026-06-27T06:32:37.255041Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Hmbown 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:32:37.255041Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MiniMax CLI is an unofficial public terminal UI/CLI for the MiniMax platform. The README states that users can chat with MiniMax-M2.5, run an approval-gated tool-using agent, and generate MiniMax media from the terminal…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public installable CLI/TUI project for terminal chat, in-app help, diagnostics, and approval-gated agent workflows against MiniMax platform APIs.

模型作用:MiniMax-M2 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:The MiniMax-M2 model line is the default text/chat model for the CLI experience, driving conversational responses and the tool-using agent loop before media-generation calls are invoked.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The artifact binds to the MiniMax-M2 family via MiniMax-M2.5 rather than the MiniMax-M2.7 variant; repository was reachable during collection.

原始记录:MiniMax CLI uses the MiniMax-M2 line for terminal chat and tool-using agents

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kevin1 使用 o3 处理软件工程任务执行

kevin1 · o3

A
厂商:OpenAI 模型:o3 来源平台:GitHub 最后复核:2026-06-27T06:28:04Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kevin1 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:28:04Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Cloudflare Workers and Twilio chatbot that answers weather forecast and general-knowledge questions over SMS, aimed at low-connectivity scenarios such as satellite messaging.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository implements the chatbot and states it is built with OpenAI o3, Cloudflare Workers, the National Weather Service API, and Twilio; the TypeScript source wires OpenAI into the SMS workflow.

模型作用:o3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o3 provides the reasoning and natural-language response generation for SMS users while the surrounding worker fetches weather data and delivers compact replies through Twilio.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The public README is concise, but it names the concrete application, technology stack, and OpenAI o3 as a component; the repository was reachable during collection.

原始记录:Backcountry AI Chat uses OpenAI o3 for SMS weather and general-knowledge assistance

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ai-asa 使用 o3 处理软件工程任务执行

ai-asa · o3

A
厂商:OpenAI 模型:o3 来源平台:GitHub 最后复核:2026-06-27T06:28:04Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ai-asa 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:28:04Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Model Context Protocol server for using OpenAI o3 directly from Claude Desktop or Claude Code, including general chat and an error-resolution tool with automatic file reading support.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository and package metadata define @ai-asa/openai-o3-mcp, document npx installation, and expose tools such as openai_chat and openai_error_resolver for o3-powered conversations and troubleshooting.

模型作用:o3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o3 is the external reasoning model used by the MCP server to answer questions, generate text, and analyze errors supplied by Claude users and coding workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The npm page returned a 403 to the collection environment, so the GitHub repository is used as the reachable artifact; README and package metadata both identify the o3 integration.

原始记录:ai-asa openai-o3-mcp connects Claude Desktop and Claude Code to OpenAI o3

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MiniMax 使用 MiniMax-M2.1 处理软件工程任务执行

MiniMax · MiniMax-M2.1

A
厂商:MiniMax 模型:MiniMax-M2.1 来源平台:official_web 最后复核:2026-06-27T06:24:42Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MiniMax 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T06:24:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MiniMax's official MiniMax-M2.1 README states that the public MiniMax Agent product is built on MiniMax-M2.1 and positions it for coding, tool use, instruction following, long-horizon planning, multilingual software dev…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A publicly accessible MiniMax Agent product is available as the user-facing artifact for running autonomous agent workflows on the model.

模型作用:MiniMax-M2.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.1 provides the agentic reasoning, tool-use, coding, and long-horizon planning capabilities that power MiniMax Agent.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Vendor-operated product rather than third-party customer story, but the evidence explicitly binds MiniMax Agent to MiniMax-M2.1 and the artifact is public.

原始记录:MiniMax Agent built on MiniMax-M2.1 for autonomous application workflows

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

fsgeek / yanantin 使用 MiniMax-M2.1 处理软件工程任务执行

fsgeek / yanantin · MiniMax-M2.1

A
厂商:MiniMax 模型:MiniMax-M2.1 来源平台:github 最后复核:2026-06-27T06:24:42Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

fsgeek / yanantin 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Yanantin repository contains a Chasqui Scout tensor report whose header records Run 3042 using model minimax/minimax-m2.1, then describes the model exploring the yanantin codebase, reading docs/cairn, src/yanantin, …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The model generated a committed markdown scout report with observations about the project's story, code structure, tensor composition, and implementation gaps.

模型作用:MiniMax-M2.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.1 performed the long-context repository inspection and synthesized the written scout observations captured in the artifact.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a generated repository artifact rather than a marketing case study; however it records the exact model, task, usage metadata, timestamp, and public output.

原始记录:Yanantin Chasqui Scout used MiniMax-M2.1 to inspect a codebase and produce an epistemic observability report

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

talkcody 使用 MiniMax-M2.1 处理代码审查和测试生成

talkcody · MiniMax-M2.1

A
厂商:MiniMax 模型:MiniMax-M2.1 来源平台:github 最后复核:2026-06-27T06:24:42Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

talkcody 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T06:24:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:TalkCody's committed API recording identifies provider MiniMax, protocol anthropic, model MiniMax-M2.1, and a user prompt asking: 'What is the relationship between Starrrocks' pipeline and fragments?' The request includ…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The recorded streamed response explains that FragmentContext contains multiple Pipelines, pipelines are chains of operator factories creating PipelineDrivers, and the fragment context coordinates lifecycle and execution…

模型作用:MiniMax-M2.1 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.1 handled the code-search-assisted reasoning and generated the explanatory answer over StarRocks internals.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a test/recording artifact inside a developer tool repository, not a customer story, but it is a concrete public trace of exact-model use with prompt, tool calls, response, and status 200.

原始记录:TalkCody recorded MiniMax-M2.1 answering a StarRocks codebase question with tool-assisted code search

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

HKUDS / DeepCode 使用 GLM-5.1 处理软件工程任务执行

HKUDS / DeepCode · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:github 最后复核:2026-06-27T09:22:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

HKUDS / DeepCode 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DeepCode is an open agentic coding system for Paper2Code, Text2Web and Text2Backend workflows; its UI can fetch OpenRouter model metadata and lets users enter exact IDs such as z-ai/glm-5.1 for default, planning and imp…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes the DeepCode application and documents runtime model selection plus configuration reloads so newly started workflows can use z-ai/glm-5.1 without manual JSON editing.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 is an explicitly documented model option for DeepCode's planning and implementation agent phases, contributing LLM reasoning and code synthesis to convert papers or text specifications into code/web/backend arti…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public DeepCode README and exact OpenRouter model id; it is a real open-source artifact, but not a single named production customer deployment.

原始记录:HKUDS DeepCode exposes z-ai/glm-5.1 as an OpenRouter model for Paper2Code and Text2Web workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

tobiasbischoff / openclawnews 使用 MiniMax-M2.1 处理软件工程任务执行

tobiasbischoff / openclawnews · MiniMax-M2.1

A
厂商:MiniMax 模型:MiniMax-M2.1 来源平台:github 最后复核:2026-06-27T06:24:42Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tobiasbischoff / openclawnews 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:24:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The OpenClawNews server code defines a public GitHub activity/news summarizer that fetches commits and issues from a configured repository, stores them in SQLite, and sets AI_MODEL to process.env.ANTHROPIC_MODEL or the …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains a runnable web app/server artifact for producing AI summaries of GitHub repository activity, with MiniMax-M2.1 selected as the default model when an API key is present.

模型作用:MiniMax-M2.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.1 supplies the AI summarization layer for transforming commit and issue data into digestible repository news summaries.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is source code at a specific commit rather than a hosted demo; the repository is public and the exact model is the default in the application code.

原始记录:OpenClawNews uses MiniMax-M2.1 as the default model for GitHub commit and issue news summaries

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Distyl 使用 GPT-4o (Aug) 处理软件工程任务执行

Distyl · GPT-4o (Aug)

A
厂商:OpenAI 模型:GPT-4o (Aug) 来源平台:official_web 最后复核:2026-06-27T06:31:07Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Distyl 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T06:31:07Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Distyl applied a fine-tuned GPT-4o model to text-to-SQL workflows including query reformulation, intent classification, chain-of-thought/self-correction, and SQL generation for enterprise data tasks.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:OpenAI reports Distyl's fine-tuned GPT-4o reached 71.83% execution accuracy and ranked first on the BIRD-SQL leaderboard at the time of the post.

模型作用:GPT-4o (Aug) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Fine-tuned GPT-4o supplied the domain-specific natural-language-to-SQL reasoning and generation capability used to improve execution accuracy across SQL tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The public evidence is an OpenAI partner success paragraph rather than a Distyl implementation deep dive. The page explicitly ties GPT-4o fine-tuning to selecting gpt-4o-2024-08-06 as the base model.

原始记录:Distyl uses fine-tuned GPT-4o for enterprise text-to-SQL generation

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Vidhyadhara Rao Kotagiri / KL Unive… 使用 Qwen2.5 72B 处理知识检索和问答

Vidhyadhara Rao Kotagiri / KL University RAG-based Assistant · Qwen2.5 72B

A
厂商:Qwen / Alibaba 模型:Qwen2.5 72B 来源平台:github 最后复核:2026-06-27T06:41:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Vidhyadhara Rao Kotagiri / KL University RAG-based Assistant 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T06:41:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Streamlit RAG assistant for answering questions related to KL University: it embeds user queries, retrieves relevant chunks from a local text corpus, and sends the retrieved context to Qwen/Qwen2.5-72B-Instruct via th…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The application returns campus-policy answers in a web UI; the README example asks about KL University dress-code policy and describes retrieving the top two similar chunks before displaying the generated answer.

模型作用:Qwen2.5 72B 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2.5-72B-Instruct is the text-generation component that turns retrieved KL University context into the final natural-language answer.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public repo evidence binds the exact Hugging Face model id; artifact is the source repository rather than a hosted production deployment.

原始记录:KL University RAG-based Assistant uses Qwen2.5-72B-Instruct for campus Q&A

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AK47-tech656 使用 Qwen2.5 72B 处理软件工程任务执行

AK47-tech656 · Qwen2.5 72B

A
厂商:Qwen / Alibaba 模型:Qwen2.5 72B 来源平台:github 最后复核:2026-06-27T06:41:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AK47-tech656 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:41:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Gradio email triage app sends sender, subject, and body fields to Qwen/Qwen2.5-72B-Instruct with business rules for spam, billing/legal threats, infrastructure incidents, and feature requests.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The app returns a structured routing decision with assigned department, priority level, and reasoning; README examples report perfect or high reward scores on phishing, refund/legal, and feature-request cases.

模型作用:Qwen2.5 72B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2.5-72B-Instruct performs the classification and JSON generation that maps each incoming email to department and priority.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:README text labels the brain as Llama-3.3-70B, but the executable app.py evidence binds the live InferenceClient call to Qwen/Qwen2.5-72B-Instruct; hosted Space referenced by README returned unauthorized during probe, s…

原始记录:Autonomous Email Triage Agent uses Qwen2.5-72B-Instruct to route support emails

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

okupacolossal / Goncalo-N 使用 Qwen2.5 72B 处理知识检索和问答

okupacolossal / Goncalo-N · Qwen2.5 72B

A
厂商:Qwen / Alibaba 模型:Qwen2.5 72B 来源平台:github 最后复核:2026-06-27T06:41:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

okupacolossal / Goncalo-N 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T06:41:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A terminal conversational assistant powered by Qwen2.5-72B-Instruct via Hugging Face Serverless Inference API implements a custom tool-calling pipeline for web search, calculation, file reads/writes, and Python executio…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The assistant maintains multi-turn history, has the model emit structured JSON tool calls, dispatches the requested tool, reinjects the result, and returns a final rich-formatted natural-language response in the termina…

模型作用:Qwen2.5 72B 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2.5-72B-Instruct drives both the conversational response and the structured JSON tool-call decisions in the agent loop.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public README directly names Qwen2.5-72B-Instruct and describes the artifact architecture; no hosted demo was required because the repository itself is the runnable artifact.

原始记录:Terminal AI assistant uses Qwen2.5-72B-Instruct for manual tool calling

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SM649 使用 Qwen2.5 Max 处理真实任务执行

SM649 · Qwen2.5 Max

A
厂商:Qwen / Alibaba 模型:Qwen2.5 Max 来源平台:GitHub 最后复核:2026-06-27T09:52:18Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SM649 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T09:52:18Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:SM649 构建一个名为 Clothesline 的 AI-generated clothing brand website,技术栈包括 HTML、CSS、JavaScript、Bootstrap 与 Flask 路由;README 明确说明前端和图片由 Google 2.5 Flash 与 Qwen2.5-Max 创建。

公开产物:公开 GitHub 仓库提供了 Clothesline 网站项目源码,仓库描述和 README 均说明这是一个现代、响应式、无数据库的服装品牌网站项目。

模型作用:Qwen2.5-Max 被明确列为生成该网站前端和图片素材的模型之一,贡献于页面实现与视觉/内容资产生成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README 与仓库描述;未给出 Qwen2.5-Max 与 Google 2.5 Flash 各自贡献比例,因此记录为联合辅助生成案例。

原始记录:SM649 用 Qwen2.5-Max 辅助生成 Clothesline 服装品牌网站前端与图像素材

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xescodes 使用 Qwen2.5 Max 处理软件工程任务执行

xescodes · Qwen2.5 Max

A
厂商:Qwen / Alibaba 模型:Qwen2.5 Max 来源平台:github 最后复核:2026-06-27T06:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xescodes 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The user built Challigrai, a Flask web application that lets writers draft text, choose a tone, and request AI-generated continuation/inspiration.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public repository containing the Challigrai app with a responsive editor, file open/save controls, tone selector, generated-text area, and Qwen API integration.

模型作用:Qwen2.5 Max 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README says the text generation capabilities are powered by Qwen 2.5 Max, and the app code calls Alibaba DashScope Generation to continue user text in a selected tone.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The README names Qwen 2.5 Max; the current app.py uses the DashScope qwen-max model alias, so exact backend version may depend on Alibaba's alias mapping at runtime.

原始记录:Challigrai uses Qwen 2.5 Max API for assisted creative writing

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Jelima64 使用 Qwen2.5 Max 处理可玩交互原型构建

Jelima64 · Qwen2.5 Max

A
厂商:Qwen / Alibaba 模型:Qwen2.5 Max 来源平台:github 最后复核:2026-06-27T06:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Jelima64 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T06:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The user used Qwen2.5Max to generate a Python/Colab program for analyzing Lotofácil lottery draws from an Excel file sourced from Brazil's Caixa lottery site.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public repository describing the generated Lotofácil program, its required Lotofácil.xlsx input, and its strategy for deriving subsequent games from prior draw-sum patterns and other criteria.

模型作用:Qwen2.5 Max 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly says the program was generated by the Qwen2.5Max AI, tying the model to creation of the lottery-analysis code and workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small individual project and README-centric evidence; still includes an explicit model attribution, concrete task, public repository artifact, and expected program behavior.

原始记录:Qwen2.5Max generated a Python Lotofácil lottery-analysis program

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Bluenot3 使用 Qwen2.5 Max 处理软件工程任务执行

Bluenot3 · Qwen2.5 Max

A
厂商:Qwen / Alibaba 模型:Qwen2.5 Max 来源平台:github 最后复核:2026-06-27T06:40:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Bluenot3 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:40:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The user built and published a Gradio chat application for Qwen2.5-Max, with system-prompt editing, history clearing, streaming responses, and concurrent request handling.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public GitHub repository containing the Gradio app code and README metadata for a Qwen2.5 Max demo chatbot.

模型作用:Qwen2.5 Max 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The app labels the chatbot as qwen2.5-max and calls DashScope Generation with model='qwen-max-0125', the January 2025 Qwen Max endpoint associated with Qwen2.5-Max, to generate streamed assistant responses.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The linked Hugging Face Space was not accessible during verification, so the GitHub repository is used as the public artifact; the source code and repo metadata remain reachable.

原始记录:Bluenot3 published a Gradio chat interface backed by Qwen2.5-Max

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

bisdom-cell / openclaw-model-bridge 使用 Qwen3 235B 处理软件工程任务执行

bisdom-cell / openclaw-model-bridge · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:github 最后复核:2026-06-27T06:41:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

bisdom-cell / openclaw-model-bridge 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:41:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository documents a reusable OpenClaw control-plane framework plus the author's live personal-assistant instance, including provider routing, tool governance, SLO/fallback handling, memory/RAG, scheduled jobs, an…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The README describes a working personal-assistant system with about 40 active scheduled jobs, dual-channel WhatsApp/Discord notifications, KB and multimodal search, and provider routing through an adapter and gateway.

模型作用:Qwen3 235B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B is named in the architecture as the primary text LLM behind the adapter, with other models used as fallbacks and Qwen2.5-VL used for image tasks; this binds the model to the assistant's reasoning/chat workloa…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub README maintained by the project, not an official Alibaba customer story; exact hosted endpoint and private runtime logs are not independently visible, but the artifact and architecture …

原始记录:openclaw-model-bridge uses Qwen3-235B as the primary LLM in a production personal-assistant control plane

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

nikinrussia-ui / dairix_ai 使用 Qwen3 235B 处理智能体流程编排

nikinrussia-ui / dairix_ai · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:github 最后复核:2026-06-27T06:41:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

nikinrussia-ui / dairix_ai 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T06:41:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project provides gonka_proxy.py and an install script that run a local proxy accepting OpenAI-format requests from OpenClaw or LiteLLM and forwarding them to Gonka's decentralized AI inference network.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The published artifact gives a one-command install flow, exposes a local proxy at http://127.0.0.1:8001, and includes LiteLLM/OpenClaw configuration that registers a model named GonkaOriginal for agent chat use.

模型作用:Qwen3 235B 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:The README's LiteLLM configuration explicitly sets the served model to openai/Qwen/Qwen3-235B-A22B-Instruct-2507-FP8 and labels it as Gonka Qwen3-235B, making Qwen3-235B the model used for OpenClaw agent responses throu…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public integration repository and README; it is an implementation artifact rather than an enterprise customer story. The repository targets Qwen3-235B-A22B-Instruct-2507-FP8, which is a concrete Qwen3-235B…

原始记录:gonka-openclaw connects OpenClaw/LiteLLM to Gonka's free Qwen3-235B inference through a local OpenAI-compatible proxy

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

abswn / gradient-chat-python 使用 Qwen3 235B 处理软件工程任务执行

abswn / gradient-chat-python · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:github 最后复核:2026-06-27T06:41:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

abswn / gradient-chat-python 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:41:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project implements an unofficial Python client for Gradient Chat, which uses the Parallax decentralized inference network, and lets developers send chat requests, maintain conversation context, select cluster mode/c…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact is a public installable Python package/repository with usage examples for generating chat replies and retrieving available models; the README lists Qwen3 235B among supported models and documents response f…

模型作用:Qwen3 235B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 235B is an explicitly supported selectable model in the client API, with Qwen3 noted as requiring the hybrid cluster mode; the model contributes the actual chat/reasoning generation returned by client.generate whe…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public GitHub README for an unofficial client and depends on the availability of Gradient Chat/Parallax; it is not a benchmark or tutorial-only page because the repository contains a concrete SDK ar…

原始记录:gradient-chat-python exposes Qwen3 235B chat and reasoning through a Python client for the Parallax decentralized inference network

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zwq2018 / Data-Copilot 使用 Qwen Chat 72B 处理研究分析和报告生成

zwq2018 / Data-Copilot · Qwen Chat 72B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 72B 来源平台:github 最后复核:2026-06-27T06:45:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zwq2018 / Data-Copilot 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T06:45:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Data-Copilot is an LLM-based system for Chinese stocks, funds, economic data, financial data, and news. Its README lists Qwen-72b-Chat as supported across all listed data sources, and lab_llms_call.py calls DashScope Ge…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact demonstrates autonomous querying, processing, analysis, prediction, and visualization of Chinese financial-market data, producing charts, tables, and text matched to the user's intent.

模型作用:Qwen Chat 72B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-72B-Chat serves as the chat/reasoning model that turns user financial-analysis requests into step-by-step outputs and tool-use decisions in the Data-Copilot workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the project repository and code path; no external production customer deployment is claimed.

原始记录:Data-Copilot uses Qwen-72B-Chat to autonomously query, analyze, predict, and visualize Chinese financial-market data

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ZhuJD-China / RainbowGPT 使用 Qwen Chat 72B 处理研究分析和报告生成

ZhuJD-China / RainbowGPT · Qwen Chat 72B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 72B 来源平台:github 最后复核:2026-06-27T06:45:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ZhuJD-China / RainbowGPT 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T06:45:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RainbowGPT is an AI agent project with a Stock Analysis module. The StockGPTAnalysis UI defines qwen_api_call(model='qwen-72b-chat') and sends system/user messages to DashScope Generation, then saves the Qwen response t…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact produces stock-analysis responses and saves them as per-stock qwen_response text files for the user interface workflow.

模型作用:Qwen Chat 72B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-72B-Chat is one of the LLM backends used by the stock-analysis agent to generate natural-language financial analysis from the prepared instruction and stock message context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source demo/application evidence; repository README describes broader model support including Qwen but does not provide production usage metrics.

原始记录:RainbowGPT uses qwen-72b-chat for stock-analysis agent responses

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

YiwenAI / AutoSAT 使用 Qwen Chat 72B 处理软件工程任务执行

YiwenAI / AutoSAT · Qwen Chat 72B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 72B 来源平台:github 最后复核:2026-06-27T06:45:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

YiwenAI / AutoSAT 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:45:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AutoSAT is a public repository for automatically optimizing SAT solvers via LLMs. In main_MultiAgent.py, the Qwen branch instantiates a LocalCallAPI with model_name='modelscope/qwen/Qwen-72B-Chat' for the multi-agent co…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact runs an iterative multi-agent optimization loop over SAT solver heuristics and evaluates generated solver variants using metrics such as PAR-2, solved count, and total time.

模型作用:Qwen Chat 72B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-72B-Chat is bound as the local LLM option that generates or evaluates code-improvement directions in the SAT-solver optimization loop.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Academic/open-source research artifact rather than a commercial deployment; accepted because it is a concrete application repo with code binding the exact model to a task, not just a leaderboard entry.

原始记录:AutoSAT uses Qwen-72B-Chat as a local LLM to optimize SAT solver heuristics

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

llmapp / OpenAI.mini 使用 Qwen Chat 72B 处理真实任务执行

llmapp / OpenAI.mini · Qwen Chat 72B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 72B 来源平台:github 最后复核:2026-06-27T06:45:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

llmapp / OpenAI.mini 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T06:45:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenAI.mini implements OpenAI-compatible APIs and a ChatGPT-like web frontend. Its model registry includes Qwen('Qwen/Qwen-72B-Chat', owner='Alibaba Cloud', model_args={'bf16': True}), and the frontend/backend code list…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Users can run the service, access Qwen-72B-Chat through OpenAI-style client libraries or LangChain, and chat with it in the web frontend using a model query parameter.

模型作用:Qwen Chat 72B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-72B-Chat is a registered local chat backend providing conversational completions behind the OpenAI-compatible API and browser chat UI.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence shows model integration and runnable open-source service; public usage volume is not disclosed.

原始记录:OpenAI.mini exposes Qwen-72B-Chat through an OpenAI-compatible API and ChatGPT-like web frontend

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hexdocom / Lemon AI 使用 Qwen Chat 72B 处理研究分析和报告生成

hexdocom / Lemon AI · Qwen Chat 72B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 72B 来源平台:github 最后复核:2026-06-27T06:45:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hexdocom / Lemon AI 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T06:45:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Lemon AI is a full-stack open-source self-evolving general AI agent for deep research, web browsing, vibe coding, data analysis, and document processing. Its completion configuration includes a provider named 'Qwen-72B-…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact provides an agent framework with planning, action, reflection, memory, code-interpreter sandbox, and configurable LLM backends including the Qwen-72B-Chat Int4 variant.

模型作用:Qwen Chat 72B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:The Qwen-72B-Chat Int4 backend is configured as one of the chat-completion providers that can power Lemon AI's agent planning and task-execution conversations.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence binds the quantized Qwen-72B-Chat-Int4 variant rather than the full-precision checkpoint; included as a deployable derivative of Qwen-72B-Chat, with no usage metrics claimed.

原始记录:Lemon AI configures a Qwen-72B-Chat-Int4 backend for local agent workflows

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

FanaHOVA 使用 o1 处理软件工程任务执行

FanaHOVA · o1

A
厂商:OpenAI 模型:o1 来源平台:github 最后复核:2026-06-27T07:08:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

FanaHOVA 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:08:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project adds code-interpreter-style execution to OpenAI o1: a user passes an analytical prompt, o1 plans the work, E2B executes generated Python, and outputs are saved locally.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The README describes generating a report on the evolution of labor productivity in the United States and displaying charts as outputs, with generated files saved to an o1_outputs folder.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The code calls the OpenAI API with the o1-preview member of the o1 family for the main reasoning/planning step, then uses auxiliary code generation/execution to materialize the analysis and charts.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository code defaults to o1-preview, so the evidence binds to the o1 family rather than only the final o1-2024-12-17 snapshot. Public repo and artifact are reachable.

原始记录:OpenAI O1 Code Interpreter wraps o1 with E2B to create reports and charts

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Maxine Xiong / OpenAI-API-Web-Apps 使用 o1 处理软件工程任务执行

Maxine Xiong / OpenAI-API-Web-Apps · o1

A
厂商:OpenAI 模型:o1 来源平台:github 最后复核:2026-06-27T07:08:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Maxine Xiong / OpenAI-API-Web-Apps 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:08:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A deployed Streamlit application named Talk to GPT lets users converse with OpenAI models through text or speech input and select o1 among the supported GPT models.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents the application, usage flow, screenshots/GIF demos, and a public Streamlit deployment link for interactive chatbot responses and optional audio playback.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI o1 is one of the selectable models powering high-quality responses in the chatbot, alongside speech-to-text and text-to-speech components for voice interaction.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The Streamlit runtime redirected through Streamlit authentication during automated HEAD checks, so the stable artifact URL is the public GitHub repository and README rather than relying on live app availability.

原始记录:OpenAI API Web Apps exposes o1 in a Streamlit Talk to GPT chatbot

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Maxine Xiong 使用 o1 处理软件工程任务执行

Maxine Xiong · o1

A
厂商:OpenAI 模型:o1 来源平台:GitHub 最后复核:2026-06-27T06:34:20Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Maxine Xiong 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:34:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a Python/Tkinter desktop application for natural-language ChatGPT conversations from a local computer, with model selection that includes OpenAI o1.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public desktop app repository with README evidence, demo video link, API-key setup, conversation history, model picker, and audio playback features.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o1 is exposed through the app as an advanced OpenAI model option for producing responses to user messages in desktop chat sessions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The app is an unofficial client and supports several OpenAI models; evidence confirms o1 is a selectable model but not that every demo interaction used o1.

原始记录:Maxine Xiong built a local ChatGPT desktop app with selectable o1

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CookSleep 使用 o1 处理软件工程任务执行

CookSleep · o1

A
厂商:OpenAI 模型:o1 来源平台:GitHub 最后复核:2026-06-27T06:34:20Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CookSleep 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:34:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Develop a PyQt5 GUI tool that removes metadata from images, including support for local and network images, concurrent processing, JPEG/WEBP metadata handling, PNG chunk filtering, and optional alpha-channel removal.

公开产物:A public GitHub repository for 图片元数据消除器 / ImageMetadataRemover with screenshot, feature list, and runnable application code.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states that the project code was mainly written by OpenAI o1-preview and OpenAI o1-mini, with the author providing feature-design proposals and feedback.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The code was co-authored with other models as well as human feedback; evidence still explicitly identifies o1-preview and o1-mini as primary contributors to the artifact.

原始记录:CookSleep used o1-preview and o1-mini to write an image metadata remover

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Cole Murray 使用 o1 处理软件工程任务执行

Cole Murray · o1

A
厂商:OpenAI 模型:o1 来源平台:GitHub 最后复核:2026-06-27T06:34:20Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Cole Murray 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:34:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Generate a TypeScript/Node implementation of OpenAI's Swarm repository with agent management, function calls, streaming responses, debugging, and example demo-loop usage.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public SwarmJS-Node repository with TypeScript library code, CI badge, quick-start instructions, and example agent workflows.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly says SwarmJS-Node is a TypeScript library implementation that o1-mini generated from OpenAI's Swarm repository.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The model variant is o1-mini, bound to the o1 family; artifact is an implementation generated from another open-source repo rather than an end-customer deployment.

原始记录:Cole Murray used o1-mini to generate a TypeScript Swarm implementation

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SherlockHao 使用 Qwen3 Max 处理多模态内容处理

SherlockHao · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:github 最后复核:2026-06-27T06:53:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SherlockHao 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T06:53:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:将 DingTalk A1 / 钉钉 AI 听记的全天录音转写与说话人分离结果输入 Qwen3-Max 262K,上下文内单次或分片生成结构化日报,覆盖决策、行动项和关键人动态。

公开产物:仓库文档给出完整架构、pipeline、成本与输出路径:短日 1 次调用、长日 3-4 次调用,最终产物为 30 秒/3 分钟/完整版结构化日报。

模型作用:Qwen3-Max 的 262,144 token 上下文被用作核心摘要生成层,支持约 90% 工作日全量文本单次理解,减少分片导致的信息割裂。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:公开仓库主要是设计与实现文档,未包含真实企业录音数据;证据明确绑定 Qwen3-Max 与具体应用任务。

原始记录:SherlockHao 用 Qwen3-Max 262K 构建全天录音智能摘要系统

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lukun927 使用 Qwen3 Max 处理知识检索和问答

lukun927 · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:github 最后复核:2026-06-27T06:53:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

lukun927 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T06:53:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:基于课程资料构建本地向量库,并通过 Streamlit Web UI 或命令行提供带来源的课程问答与段落改写能力。

公开产物:公开仓库包含可运行的 app.py、main.py、RAGAgent、文档加载、切分、ChromaDB 向量库和多模态工具,README 标注运行方式;config.py 将 MODEL_NAME 固定为 qwen3-max。

模型作用:Qwen3-Max 作为 RAG 代理的生成模型,结合检索结果回答课程问题,并支持对回答段落按用户指令进行重写。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目为公开课程助教应用仓库,未披露具体课程数据;模型名、任务和可运行产物均可从仓库核验。

原始记录:lukun927 用 Qwen3-Max 和本地 RAG 实现课程助教 Agent

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mesakurax 使用 Qwen3 Max 处理软件工程任务执行

mesakurax · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:github 最后复核:2026-06-27T06:53:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mesakurax 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:53:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:针对 Hotel Pricing System 架构作业,用直接 LLM 交互范式让 Qwen3-Max 按 ADD 3.0 方法完成四轮架构设计迭代。

公开产物:仓库提供 CLI runner、PDF 作业材料、prompt、测试、report.md 模板,并在真实运行后生成 conversation_log、add_results 与 report_notes 等交付物。

模型作用:Qwen3-Max 通过 DashScope/OpenAI-compatible API 接收包含 ADD 3.0 方法、案例研究和迭代计划的系统提示,逐轮产出架构决策、视图代码和报告素材。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:案例是课程/作业型工程产物,不是生产系统;但公开仓库明确记录模型、使用者、任务、运行方式和产出文件。

原始记录:mesakurax 用 Qwen3-Max 生成酒店定价系统 ADD 3.0 架构设计交付物

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

qihaozh 使用 Qwen3 Max 处理真实任务执行

qihaozh · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:github 最后复核:2026-06-27T06:53:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

qihaozh 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T06:53:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:用三个专门 AI 顾问角色(战略、财务、运营)围绕用户输入的创业想法协作生成完整商业计划。

公开产物:公开仓库提供 CLI 应用,可输出 business_plan.json 和 business_plan.html;README 展示针对 AI fitness app、sustainable fashion marketplace、EdTech platform 等输入生成 JSON/HTML 商业计划的命令。

模型作用:Qwen3-Max 驱动多智能体对话与内容生成,负责市场分析、竞争定位、财务计划、运营计划等结构化商业计划内容。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:公开证据来自项目 README 和源码,未展示真实商业客户;作为开源产品/应用原型案例可核验。

原始记录:qihaozh 用 Qwen3-Max 构建多智能体创业商业计划生成器

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

git-bob / Robert Haase 使用 GPT-4o (Aug) 处理软件工程任务执行

git-bob / Robert Haase · GPT-4o (Aug)

A
厂商:OpenAI 模型:GPT-4o (Aug) 来源平台:github 最后复核:2026-06-27T06:53:58Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

git-bob / Robert Haase 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:53:58Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:git-bob is an AI agent that runs in GitHub CI/GitLab runners to understand repository issues and pull requests, comment on issues, review PRs, split issues, and generate code or text changes for software-development wor…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository documents the working tool, package, CI workflow usage, and online issue-discussion artifacts; its README lists openai:gpt-4o-2024-08-06 among tested/configurable LLMs and states that tested GPT-4o…

模型作用:GPT-4o (Aug) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o 2024-08-06 supplies the language understanding and code/text generation used by git-bob to interpret GitHub issues or pull requests and produce implementation, review, or comment outputs inside CI.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source tool evidence; per-run model choice is user-configured via GIT_BOB_LLM_NAME, so claim support/tested use rather than exclusive default production routing.

原始记录:git-bob uses GPT-4o 2024-08-06 to solve GitHub issues and pull requests

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ai-okx 使用 Qwen3 Max 处理金融和商业分析

ai-okx · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:github 最后复核:2026-06-27T06:53:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ai-okx 公开的金融与商业分析案例,来源为 公开代码库,复核于 2026-06-27T06:53:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建 BTC/USDT 加密货币 AI 自动交易机器人,接入 OKX 交易所,周期性让 AI 模型根据市场数据输出 buy/sell/hold 决策并设置止盈止损。

公开产物:仓库包含交易程序、Streamlit Web 监控界面、账户/持仓/收益曲线/AI 决策展示、止盈止损订单管理等文件;仓库描述明确支持 deepseek 或 qwen3-max 自动炒币。

模型作用:qwen3-max 被列为交易决策模型选项,用于分析行情并生成交易信号、信心和仓位相关建议,机器人再执行交易与风险控制逻辑。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 当前默认说明偏 DeepSeek,qwen3-max 绑定主要来自 GitHub 仓库描述和配置模板的 Qwen/DashScope 支持;项目公开可访问但外部演示域名核验时不可解析,因此 artifact 使用 GitHub 仓库。

原始记录:ai-okx 在 Alpha Arena OKX 自动交易机器人中提供 qwen3-max 决策选项

已有真实案例 金融与商业分析公开代码库A 类可核验real_case auto_approved 进入模型卡精选

uni-medical / MedSegAgent authors 使用 GPT-4o (Aug) 处理软件工程任务执行

uni-medical / MedSegAgent authors · GPT-4o (Aug)

A
厂商:OpenAI 模型:GPT-4o (Aug) 来源平台:github 最后复核:2026-06-27T06:53:58Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

uni-medical / MedSegAgent authors 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:53:58Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MedSegAgent is a multi-agent system for instructive medical image segmentation: it parses free-form clinical segmentation requests, filters candidate datasets from modality to anatomy to label, and selects/runs speciali…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains the accepted JBHI project code, dataset metadata, evaluation script, and OAI_CONFIG_LIST example that explicitly configures model gpt-4o-2024-08-06 for the OpenAI tag; eval_example.sh also invoke…

模型作用:GPT-4o (Aug) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o 2024-08-06 is used as the LLM component for natural-language request understanding and coarse-to-fine segmentation model selection before medical segmentation execution.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Healthcare research artifact rather than a deployed clinical product; do not imply clinical approval or live patient use.

原始记录:MedSegAgent uses GPT-4o 2024-08-06 for instructive medical image segmentation model selection

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Simon Schubert / Kai 9000 使用 GPT-4o (Aug) 处理软件工程任务执行

Simon Schubert / Kai 9000 · GPT-4o (Aug)

A
厂商:OpenAI 模型:GPT-4o (Aug) 来源平台:github 最后复核:2026-06-27T06:53:58Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Simon Schubert / Kai 9000 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:53:58Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Kai 9000 is an open-source AI assistant with persistent memory across Android, iOS, Windows, macOS, Linux, and Web; it supports chat, generated interactive screens, reminders, email/memory background checks, and tool/MC…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository and public web app are reachable; the model catalog explicitly maps gpt-4o-2024-08-06 to GPT-4o with a 128k context window and the README lists OpenAI as a supported service for the assistant product.

模型作用:GPT-4o (Aug) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o 2024-08-06 is available as an OpenAI model option that can provide the assistant's conversation, reasoning, screen-generation, and memory/task-handling responses when selected by the user/provider configuration.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence confirms product support and exact model catalog entry; it does not prove every Kai deployment defaults to GPT-4o 2024-08-06.

原始记录:Kai 9000 exposes GPT-4o 2024-08-06 in a cross-platform persistent-memory AI assistant

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

HKUDS / AI-Researcher 使用 GPT-4o (Aug) 处理软件工程任务执行

HKUDS / AI-Researcher · GPT-4o (Aug)

A
厂商:OpenAI 模型:GPT-4o (Aug) 来源平台:github 最后复核:2026-06-27T06:53:58Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

HKUDS / AI-Researcher 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:53:58Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AI-Researcher is an autonomous scientific discovery system that helps researchers move from reference papers or research prompts through idea generation, experiment planning/execution, and paper-writing workflows.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project page, documentation, and repository are public; the research_agent constant file sets COMPLETION_MODEL from the environment with gpt-4o-2024-08-06 as the fallback value used by the agent code path.

模型作用:GPT-4o (Aug) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-4o 2024-08-06 acts as the completion model for the research agent's planning, idea generation, tool orchestration, and research-writing steps when the default or OpenAI configuration is used.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Project documentation also shows other configurable providers; treat this as an open-source agent integration/default in code, not proof that all public demos currently route to GPT-4o 2024-08-06.

原始记录:AI-Researcher includes GPT-4o 2024-08-06 as the completion model for autonomous scientific research workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

opencode-ai 使用 QwQ-32B 处理软件工程任务执行

opencode-ai · QwQ-32B

A
厂商:Qwen / Alibaba 模型:QwQ-32B 来源平台:github 最后复核:2026-06-27T06:54:55Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

opencode-ai 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T06:54:55Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:opencode is a terminal coding agent; its Groq model registry binds the QWEN Qwq model to API model qwen-qwq-32b, and the configuration defaults the coder, summarizer, task, and title agents to that model.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public coding-agent repository that can run code-editing, task execution, summarization, and title-generation workflows with QwQ-32B via Groq.

模型作用:QwQ-32B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:QwQ-32B supplies the reasoning LLM behind the default Groq-backed coding agent roles, contributing code/task reasoning and summarization behavior inside the tool.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository source code and defaults; no independent production usage metrics were found in this pass.

原始记录:opencode uses QwQ-32B as its default Groq model for coding-agent workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AstraBert 使用 QwQ-32B 处理知识检索和问答

AstraBert · QwQ-32B

A
厂商:Qwen / Alibaba 模型:QwQ-32B 来源平台:github 最后复核:2026-06-27T06:54:55Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AstraBert 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T06:54:55Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RAGcoon is an agentic RAG application for helping users build a startup. Its README states that the main workflow is handled by a ReAct Query Agent using QwQ-32B provisioned by Groq, and the implementation instantiates …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A runnable Docker/source application with frontend, backend, Qdrant retrieval, and a QwQ-32B-based Query Agent for answering startup-building questions from retrieved context.

模型作用:QwQ-32B 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:QwQ-32B acts as the base reasoning model for the ReAct query agent, deciding how to use retrieved information and generate the final user-facing answer.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is direct repository documentation and source code; case is an open-source artifact rather than a separately published customer story.

原始记录:RAGcoon uses QwQ-32B as the ReAct query agent for startup-building RAG workflows

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

wassim249 使用 QwQ-32B 处理知识检索和问答

wassim249 · QwQ-32B

A
厂商:Qwen / Alibaba 模型:QwQ-32B 来源平台:github 最后复核:2026-06-27T06:54:55Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

wassim249 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T06:54:55Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:YT Navigator is an AI-powered application for scanning YouTube channel videos, indexing transcripts, semantic search, and chatting with channel content. Its README identifies qwen-qwq-32b and llama-3.1-8b-instant from G…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public application that processes up to 100 channel videos, stores transcript segments in PGVector, and lets users search or ask questions with timestamped answers grounded in channel content.

模型作用:QwQ-32B 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:QwQ-32B serves as the powerful reasoning LLM in the ReAct agent path, helping interpret user questions, reason over retrieved transcript context, and produce chat/search responses.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the repository README and configuration defaults; live hosted deployment availability was not required or confirmed.

原始记录:YT Navigator uses QwQ-32B for YouTube-channel search and chat agents

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

cRED-f 使用 QwQ-32B 处理研究分析和报告生成

cRED-f · QwQ-32B

A
厂商:Qwen / Alibaba 模型:QwQ-32B 来源平台:github 最后复核:2026-06-27T06:54:55Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

cRED-f 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T06:54:55Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:QuestGen-AI is an exam question generation platform where users upload PDF content and configure generation; its README states the default AI model is qwen/qwq-32b:free, and the API route falls back to that same model w…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public web application that generates customized question papers from uploaded educational PDF material using an OpenRouter model selection flow.

模型作用:QwQ-32B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:QwQ-32B is the default generation model, contributing reasoning and language generation for converting source PDF content into exam-style questions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is direct repository documentation and source code; public repo does not provide independent user-volume or deployment metrics.

原始记录:QuestGen AI Agent uses QwQ-32B to generate customized exam questions from uploaded PDFs

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

enricollen 使用 QwQ-32B 处理文档理解和结构化处理

enricollen · QwQ-32B

A
厂商:Qwen / Alibaba 模型:QwQ-32B 来源平台:github 最后复核:2026-06-27T06:54:55Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

enricollen 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T06:54:55Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:fastRTC Voice Agent is a realtime voice-enabled AI assistant. Its setup documentation sets OPENROUTER_MODEL=qwen/qwq-32b:free, and the LLM service defaults OpenRouter to qwen/qwq-32b:free.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public voice-agent application that can run realtime spoken conversations through a web/RTC interface and route LLM responses through OpenRouter.

模型作用:QwQ-32B 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:QwQ-32B provides the language reasoning and response generation component of the voice assistant when using the OpenRouter backend.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from public repository docs and code defaults; hosted demo uptime or active-user claims were not independently verified.

原始记录:fastRTC Voice Agent uses QwQ-32B as its OpenRouter LLM for real-time voice conversations

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

coreline-ai 使用 GLM-4.5 处理知识检索和问答

coreline-ai · GLM-4.5

A
厂商:Z AI / GLM 模型:GLM-4.5 来源平台:github 最后复核:2026-06-27T07:00:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

coreline-ai 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T07:00:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:coreline-ai 发布 Antigravity GLM MCP,把 Gemini/Antigravity Agent 通过 MCP 连接到云端 GLM-4.5 API,用于把复杂编码任务委派给 GLM,并提供文件、Git、Web Search、记忆和备份等 25 个自动化工具。

公开产物:公开仓库提供可安装的 MCP Server、架构文档、快速开始和工具参考;README 的架构图明确显示 MCP Server 调用 GLM-4.5 API,配置示例将 GLM_MODEL 设置为 GLM-4.5。

模型作用:GLM-4.5 是该工具的核心推理/编码后端,负责接收 Antigravity 代理委派的复杂编码请求并结合 MCP 工具完成自动化开发任务。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;未验证实际生产部署规模,但仓库本身是可访问的公开工程产物且明确绑定 GLM-4.5。

原始记录:Antigravity GLM MCP 将 Gemini/Antigravity 编码代理桥接到 GLM-4.5

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

fakerybakery 使用 GLM-4.5 处理软件工程任务执行

fakerybakery · GLM-4.5

A
厂商:Z AI / GLM 模型:GLM-4.5 来源平台:github 最后复核:2026-06-27T07:00:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

fakerybakery 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:00:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:fakerybakery 的 OpenBridge 是一个开源 API bridge,用于让 GLM-4.5 等 OpenAI-compatible LLM 在 Claude Code 中以 Anthropic-compatible tools 工作。

公开产物:公开仓库和 README 提供 pip 安装、openbridge 代理启动方式,以及 OPENAI_MODEL="zai-org/GLM-4.5:fireworks-ai" obcli 的 GLM-4.5 使用示例。

模型作用:GLM-4.5 在该案例中作为 Claude Code 背后的可替换 LLM,承担代码代理会话中的工具调用、生成和推理任务。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;该项目支持多模型,GLM-4.5 是 README 明确列出的可用模型之一。

原始记录:OpenBridge 让 GLM-4.5 接入 Claude Code 的 Anthropic 兼容工具链

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xxoo6699 使用 GLM-4.5 处理真实任务执行

xxoo6699 · GLM-4.5

A
厂商:Z AI / GLM 模型:GLM-4.5 来源平台:github 最后复核:2026-06-27T07:00:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xxoo6699 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:00:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:xxoo6699 的 ZtoApi 为 Z.ai GLM-4.5 构建 OpenAI 兼容 API 代理服务器,支持标准 OpenAI API 请求格式、流式/非流式响应、API 密钥验证、Docker/Render 部署和实时监控仪表板。

公开产物:公开仓库提供 Go 代理服务、API 文档入口、Dashboard 说明、Docker 部署方式和 MODEL_NAME 默认 GLM-4.5 的环境变量配置。

模型作用:GLM-4.5 是代理服务器转发的默认目标模型,负责实际对话、推理和流式生成;项目把其能力封装为 OpenAI-compatible 接口以接入现有客户端。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;项目是社区二次开发代理,不代表 Z.ai 官方客户案例。

原始记录:ZtoApi 为 Z.ai GLM-4.5 提供 OpenAI 兼容 API 代理和监控面板

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hmjz100 使用 GLM-4.5 处理真实任务执行

hmjz100 · GLM-4.5

A
厂商:Z AI / GLM 模型:GLM-4.5 来源平台:github 最后复核:2026-06-27T07:00:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hmjz100 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:00:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:hmjz100 发布 Z.ai2api,把 Z.ai Chat 代理为 OpenAI Compatible 格式,支持根据官网 /api/models 生成模型列表、匿名模式、图片上传和思考链格式转换;README 中 MODEL 默认值为 GLM-4.5。

公开产物:公开仓库提供 Python 服务端、环境变量配置和运行说明,允许用户用 OpenAI-compatible 客户端调用默认的 GLM-4.5。

模型作用:GLM-4.5 是该代理的默认生成模型,负责上游对话/推理输出;项目主要贡献是把 GLM-4.5 的响应与思考链转换为 OpenAI-compatible 结构。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;项目可选模型可能随 Z.ai 模型列表变化,但 README 明确声明默认模型为 GLM-4.5。

原始记录:Z.ai2api 将 Z.ai/GLM-4.5 转换为 OpenAI Compatible 服务

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

roseforyou 使用 GLM-4.5 处理真实任务执行

roseforyou · GLM-4.5

A
厂商:Z AI / GLM 模型:GLM-4.5 来源平台:github 最后复核:2026-06-27T07:00:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

roseforyou 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:00:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:roseforyou 的 ZtoApi 使用 Deno/TypeScript 为 Z.ai 的 GLM-4.5 和 GLM-4.5V 构建 OpenAI 兼容 API 代理,支持流式响应、思考过程处理、实时 Dashboard、API 密钥验证、匿名 token 和 Deno Deploy/自托管部署。

公开产物:公开仓库 README 明确列出模型 0727-360B-API = GLM-4.5,并说明 GLM-4.5 支持通用对话、代码生成、MCP 工具调用;仓库提供可部署的 main.ts 服务。

模型作用:GLM-4.5 是该代理的文本生成、代码生成和工具调用后端,代理层把模型能力转成 OpenAI API 格式并提供监控与部署能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README;同一生态内有多个 ZtoApi 实现,本条按不同作者和 Deno/TypeScript artifact 单独记录。

原始记录:Deno 版 ZtoApi 为 GLM-4.5/GLM-4.5V 提供高性能 OpenAI 兼容代理

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Edouard Foussier / Xiexie 使用 GLM-4.6 处理软件工程任务执行

Edouard Foussier / Xiexie · GLM-4.6

A
厂商:Z AI / GLM 模型:GLM-4.6 来源平台:github 最后复核:2026-06-27T07:00:02Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Edouard Foussier / Xiexie 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:00:02Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A voice-first macOS companion for seniors that reads the screen, answers spoken questions, points to UI elements, and helps identify scam emails.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public hackathon-winning repository, landing page, and 3-minute demo showing Xiexie analyzing suspicious emails and guiding the user on screen.

模型作用:GLM-4.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly states the project was built on Z.AI GLM-4.6 plus GLM-4.5V; GLM-4.6 is used as part of the agent stack for long-horizon assistant reasoning and interaction, while the visual email-reading flow uses…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:GLM-4.6 is one of multiple models/services in the stack; visual screenshot analysis is attributed to GLM-4.5V, so cite this case as a multi-model agent application rather than GLM-4.6-only vision evidence.

原始记录:Xiexie: senior scam-shield macOS companion built for GOSIM Agentic Hackathon

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yasinaky 使用 GLM-4.6 处理软件工程任务执行

yasinaky · GLM-4.6

A
厂商:Z AI / GLM 模型:GLM-4.6 来源平台:github 最后复核:2026-06-27T07:00:02Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yasinaky 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:00:02Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:An MCP server that lets Claude Code call Z.ai GLM models for subtasks, agent work, parallel processing, cross-validation, debugging, and cost optimization.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public Python MCP server repository with setup instructions, smoke tests, API-key configuration, and documented use through Claude Code MCP.

模型作用:GLM-4.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README names GLM-4.6 as the flagship model for reasoning, coding, and agentic tasks and describes offloading subtasks to GLM-4.6 from Claude Code.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The artifact is an integration tool rather than an end-user product; evidence is strong for real runtime use of GLM-4.6 in coding-agent workflows.

原始记录:GLM MCP Server for delegating Claude Code tasks to GLM-4.6

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CyPack 使用 GLM-4.6 处理软件工程任务执行

CyPack · GLM-4.6

A
厂商:Z AI / GLM 模型:GLM-4.6 来源平台:github 最后复核:2026-06-27T07:00:02Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CyPack 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:00:02Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Spanish-language MCP server that exposes a glm_route tool so Claude Code can invoke another Claude CLI instance configured against the Z.AI Anthropic-compatible endpoint.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository with MCP server code, architecture diagram text, environment configuration, and installation steps for routing prompts to GLM-4.6.

模型作用:GLM-4.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly lists the architecture path ending in 'Modelo GLM-4.6' and states glm_route routes prompts to GLM-4.6 for code generation, deep analysis, and general queries.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Runtime depends on the user's Z.AI API credentials; the public artifact demonstrates the integration and intended GLM-4.6 usage, not a hosted service.

原始记录:CCGLM MCP Server routing Claude Code prompts to Z.AI GLM-4.6

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lokafinnsw 使用 GLM-4.6 处理软件工程任务执行

lokafinnsw · GLM-4.6

A
厂商:Z AI / GLM 模型:GLM-4.6 来源平台:github 最后复核:2026-06-27T07:00:02Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

lokafinnsw 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:00:02Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Rust command-line coding assistant for code explanation, documentation, refactoring suggestions, bug detection, and interactive or single-query coding help.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public CLI repository with installation instructions, model selection, API-key setup, and documented commands for using the coding assistant.

模型作用:GLM-4.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README describes the tool as an AI-powered coding assistant using Z.ai's GLM-4.6 model and lists GLM-4.6 with a 200K context window for complex coding tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository appears to be an independent implementation and not an official Z.AI repo despite README clone text referencing z-ai/zai-coding-agent; evidence still clearly binds the artifact to GLM-4.6 usage.

原始记录:Z.ai Coding Agent CLI using GLM-4.6 for large-codebase coding tasks

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Auxiar 使用 GLM-4.6 处理软件工程任务执行

Auxiar · GLM-4.6

A
厂商:Z AI / GLM 模型:GLM-4.6 来源平台:github 最后复核:2026-06-27T07:00:02Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Auxiar 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:00:02Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A web page that calls Google's Places API to retrieve a direct 'Leave a Review' Google link from a supplied Places API key.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public GitHub repository and hosted GitHub Pages app for generating Google review links.

模型作用:GLM-4.6 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The GitHub repository description states the webpage was quickly coded with help from GLM-4.6 and Claude, explicitly recommending GLM-4.6.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from repository metadata rather than README text, and the author credits Claude alongside GLM-4.6; treat as co-assisted coding evidence.

原始记录:GoogleReviewHelper web app coded with assistance from GLM-4.6

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Dwayne Charrington / I Like Kill Ne… 使用 Gemini 1.0 Ultra 处理软件工程任务执行

Dwayne Charrington / I Like Kill Nerds · Gemini 1.0 Ultra

A
厂商:Google / Gemini 模型:Gemini 1.0 Ultra 来源平台:web_search 最后复核:2026-06-27T09:46:00Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Dwayne Charrington / I Like Kill Nerds 公开的代码代理与软件工程案例,来源为 公开网页,复核于 2026-06-27T09:46:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The author tested Gemini Advanced, explicitly described in the article title as powered by Gemini Ultra 1.0, with coding prompts for TypeScript and JavaScript generation.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The article reports that the generated code ran the first time without issue and used modern best practices close to how the author would personally write TypeScript and JavaScript.

模型作用:Gemini 1.0 Ultra 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.0 Ultra generated working source code from natural-language coding prompts and supported a long hands-on session of more than 40 prompts without the usage-cap interruption the author associated with GPT-4.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Independent hands-on practitioner article rather than a deployed public repo; artifact is the public screenshot asset embedded by the article, and the article text gives the concrete task and reported result. Verified a…

原始记录:Dwayne Charrington used Gemini Ultra 1.0 in Gemini Advanced for TypeScript and JavaScript code generation

已有真实案例 代码代理与软件工程公开网页A 类可核验real_case auto_approved 进入模型卡精选

The Verge / Emilia David 使用 Gemini 1.0 Ultra 处理多模态内容处理

The Verge / Emilia David · Gemini 1.0 Ultra

A
厂商:Google / Gemini 模型:Gemini 1.0 Ultra 来源平台:web_search 最后复核:2026-06-27T07:06:31Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

The Verge / Emilia David 公开的多模态生成与理解案例,来源为 公开网页,复核于 2026-06-27T07:06:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Verge paid for and tested Gemini Advanced, which the article states runs on Gemini Ultra, using prompts for image generation, current-events answers and Google Maps/business-information style tasks.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Gemini Ultra produced a dog image from the prompt and, in the article's assessment, performed better on Google-connected tasks such as current information and business details, while the generated dog image had visible …

模型作用:Gemini 1.0 Ultra 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.0 Ultra powered the Gemini Advanced responses, translating a detailed natural-language prompt into an image artifact and using Google ecosystem context for information-seeking tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an editorial hands-on article with mixed results, not a customer deployment; artifact URL is the public generated image embedded in the article.

原始记录:The Verge used Gemini Advanced / Gemini Ultra for image generation and Google-connected assistant tasks

已有真实案例 多模态生成与理解公开网页A 类可核验real_case auto_approved 进入模型卡精选

Understanding AI / Timothy B. Lee 使用 Gemini 1.0 Ultra 处理智能体流程编排

Understanding AI / Timothy B. Lee · Gemini 1.0 Ultra

A
厂商:Google / Gemini 模型:Gemini 1.0 Ultra 来源平台:web_search 最后复核:2026-06-27T07:06:31Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Understanding AI / Timothy B. Lee 公开的智能体工作流案例,来源为 公开网页,复核于 2026-06-27T07:06:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The author used Gemini Advanced, described in the article as the premium Gemini version powered by Gemini Ultra 1.0, as an expert copy editor on a deliberately modified draft of a prior article containing spelling and g…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Gemini identified at least one real typo ('their' that should have been 'there') but mostly gave stylistic advice, usually caught only one or two typos per run, and sometimes hallucinated nonexistent errors.

模型作用:Gemini 1.0 Ultra 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.0 Ultra performed draft review and error detection from a pasted long-form article, contributing proofreading suggestions and typo detection even though the article judged its reliability below ChatGPT in this …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Negative/mixed hands-on use case; still concrete because it identifies the user, document-editing task, source article and reported output. The article contains a likely typo naming 'Gemini Pro 1.0 Ultra', but context a…

原始记录:Understanding AI used Gemini Advanced / Gemini Ultra 1.0 to proofread a draft article

已有真实案例 智能体工作流公开网页A 类可核验real_case auto_approved 进入模型卡精选

OpenRouter 使用 Qwen3 235B 2507 处理真实任务执行

OpenRouter · Qwen3 235B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 2507 来源平台:official_web 最后复核:2026-06-27T07:29:47Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

OpenRouter 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T07:29:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenRouter lists the exact Qwen/Qwen3-235B-A22B-Instruct-2507 weights as the model behind its qwen/qwen3-235b-a22b-2507 route and makes the model available through its model page, playground and quick-start flow for gen…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public OpenRouter product artifact exists for the model, including the model slug, model-weights link, context length, pricing, provider/performance sections and playground/quick-start entry points for API users.

模型作用:Qwen3 235B 2507 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B-Instruct-2507 supplies the multilingual instruction-following, long-context reasoning, coding, math and tool-use capabilities that OpenRouter exposes to developers through its unified inference router.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public model/API product page rather than a named end-customer story; the page explicitly binds the OpenRouter route to Qwen/Qwen3-235B-A22B-Instruct-2507 and provides an accessible artifact.

原始记录:OpenRouter exposes Qwen3 235B A22B Instruct 2507 for chat-completion routing and playground use

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

DeepInfra 使用 Qwen3 235B 2507 处理真实任务执行

DeepInfra · Qwen3 235B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 2507 来源平台:official_web 最后复核:2026-06-27T07:29:47Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

DeepInfra 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T07:29:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DeepInfra publishes an API reference showing how to send chat-completion requests to its OpenAI-compatible endpoint using model "Qwen/Qwen3-235B-A22B-Instruct-2507" for conversational text generation and multi-turn prom…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public DeepInfra demo/API artifact includes a concrete POST endpoint, request payload with the exact model ID, sample conversational messages and an example completion response generated under that model name.

模型作用:Qwen3 235B 2507 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B-Instruct-2507 is the selected backend model that produces the chat-completion responses exposed by DeepInfra for application developers.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is provider documentation and a demo page, not a third-party customer narrative; it is still a concrete public deployment with exact model ID and runnable API artifact.

原始记录:DeepInfra serves Qwen3 235B A22B Instruct 2507 through an OpenAI-compatible Chat Completions API

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

Replicate 使用 Qwen3 235B 2507 处理真实任务执行

Replicate · Qwen3 235B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 2507 来源平台:official_web 最后复核:2026-06-27T07:29:47Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Replicate 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T07:29:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Replicate provides a public model page for qwen/qwen3-235b-a22b-instruct-2507 so users can run the Qwen3-235B-A22B-Instruct-2507 model through Replicate-hosted inference for text-generation tasks.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The accessible Replicate artifact presents the model as a runnable API endpoint and includes model overview details for the exact Qwen3-235B-A22B-Instruct-2507 family, linking to Qwen Chat and upstream Qwen materials.

模型作用:Qwen3 235B 2507 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B-Instruct-2507 provides the instruction-following, reasoning, coding, tool-use and 256K-class long-context language-generation capability that Replicate exposes as a hosted model.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Replicate page text mirrors upstream model-card content in parts, but the artifact is a public product endpoint for running the exact model, not a benchmark-only or collection-only page.

原始记录:Replicate packages Qwen3 235B A22B Instruct 2507 as a runnable public API model

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

Amazon Web Services / Amazon Bedrock 使用 Qwen3 235B 2507 处理软件工程任务执行

Amazon Web Services / Amazon Bedrock · Qwen3 235B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 2507 来源平台:official_web 最后复核:2026-06-27T07:29:47Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Amazon Web Services / Amazon Bedrock 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T07:29:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AWS announced Qwen3-235B-A22B-Instruct-2507 as one of the Qwen models available in Amazon Bedrock, giving customers managed serverless access through Bedrock for code generation, repository analysis, agentic workflows, …

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The AWS News Blog names Qwen3-235B-A22B-Instruct-2507 in the Bedrock launch and describes Bedrock as a unified API where customers can integrate models into applications without managing infrastructure.

模型作用:Qwen3 235B 2507 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B-Instruct-2507 contributes the large MoE instruction model capability within Bedrock's managed model catalog for long-context reasoning, coding and agentic application use cases.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an official launch/customer-access post for a managed service, not an individual workload case study; it names the exact model and concrete Bedrock application tasks.

原始记录:Amazon Bedrock adds Qwen3 235B A22B Instruct 2507 as a managed foundation model for enterprise applications

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

SiliconFlow 使用 Qwen3 235B 2507 处理研究分析和报告生成

SiliconFlow · Qwen3 235B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 2507 来源平台:official_web 最后复核:2026-06-27T07:29:47Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

SiliconFlow 公开的研究与报告生成案例,来源为 官方页面,复核于 2026-06-27T07:29:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SiliconFlow publishes a dedicated Qwen3-235B-A22B-Instruct-2507 model page with API-reference entry points and use-case sections for ultra-long document synthesis, advanced codebase analysis and other demanding text-gen…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public SiliconFlow artifact identifies Qwen3-235B-A22B-Instruct-2507 by name, describes its Alibaba Cloud Qwen origin, MoE architecture and long-context capability, and provides a path to the chat-completions API re…

模型作用:Qwen3 235B 2507 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B-Instruct-2507 supplies SiliconFlow users with the long-context understanding, instruction following, reasoning, coding and tool-use abilities highlighted on the platform's model page.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a hosted-model product page with use-case descriptions rather than a named customer deployment; the artifact is public, reachable and model-specific.

原始记录:SiliconFlow offers Qwen3 235B A22B Instruct 2507 through its model API for long-context synthesis and coding tasks

已有真实案例 研究与报告生成官方页面A 类可核验real_case auto_approved 进入模型卡精选

eosphoros-ai / DB-GPT-Hub 使用 Qwen Chat 14B 处理软件工程任务执行

eosphoros-ai / DB-GPT-Hub · Qwen Chat 14B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 14B 来源平台:github 最后复核:2026-06-27T07:30:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

eosphoros-ai / DB-GPT-Hub 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:30:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DB-GPT-Hub is a public Text-to-SQL project that supports Qwen-family base models and reports Qwen-14B-Chat LoRA/QLoRA runs for translating natural-language questions into executable SQL over database schemas.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes training/evaluation code and reported Qwen-14B-Chat Text2SQL results, including LoRA and QLoRA rows for easy, medium, hard, extra and overall execution-accuracy categories.

模型作用:Qwen Chat 14B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-14B-Chat is the base chat model adapted with LoRA/QLoRA so DB-GPT-Hub can generate SQL queries from user instructions and database context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The README includes evaluation metrics, but the artifact is a full open-source Text2SQL fine-tuning project with model-specific Qwen-14B-Chat runs, not a leaderboard-only record.

原始记录:DB-GPT-Hub fine-tunes and evaluates Qwen-14B-Chat for Text-to-SQL generation

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

junyuyang7 / ChatAgent_RAG 使用 Qwen Chat 14B 处理软件工程任务执行

junyuyang7 / ChatAgent_RAG · Qwen Chat 14B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 14B 来源平台:github 最后复核:2026-06-27T07:30:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

junyuyang7 / ChatAgent_RAG 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:30:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ChatAgent_RAG is a public project for offline deployment of large models to build a WebUI that uploads a local knowledge base for RAG question answering and is being extended toward tool-calling agents; its model config…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides runnable scripts for database initialization and WebUI startup, plus screenshots of LLM chat, knowledge-base management and knowledge-base chat pages.

模型作用:Qwen Chat 14B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-14B-Chat is one of the local LLM backends that can answer knowledge-base-grounded questions and participate in the planned agent workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The default config in the README uses another local model, but the public model registry explicitly binds Qwen-14B-Chat to the project; evidence is a real repo rather than a tutorial-only page.

原始记录:ChatAgent_RAG maps Qwen-14B-Chat into a local RAG knowledge-base WebUI

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Bart Czernicki / RiskAnalysisWithGe… 使用 o1 处理软件工程任务执行

Bart Czernicki / RiskAnalysisWithGenerativeAIReasoning · o1

A
厂商:OpenAI 模型:o1 来源平台:github 最后复核:2026-06-27T07:08:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Bart Czernicki / RiskAnalysisWithGenerativeAIReasoning 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:08:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A C# console pipeline analyzes Microsoft 2023 and 2024 SEC 10-K risk-factor sections, compares changes between filings, and consolidates material business risks into Markdown tables.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes o1-generated output files, including per-section Markdown analyses and an o1 consolidated risk analysis with risk factors, changes, potential impact, and key insights.

模型作用:o1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI o1 is the reasoning model selected for the high-token, multi-step comparison and consolidation stages; the code includes o1-specific reasoning settings and routes risk-factor prompts to the o1 deployment.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public GitHub repository with generated artifacts. Evidence is a developer/author project rather than a third-party enterprise customer story, but it contains concrete o1 outputs and code paths for the task.

原始记录:Risk Analysis with Generative AI Reasoning uses o1 to compare Microsoft SEC 10-K risk factors

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ZJUNLP / EasyInstruct 使用 Qwen Chat 14B 处理软件工程任务执行

ZJUNLP / EasyInstruct · Qwen Chat 14B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 14B 来源平台:github 最后复核:2026-06-27T07:30:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ZJUNLP / EasyInstruct 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:30:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:EasyInstruct is a public research toolkit for instruction generation, selection and prompting, including Self-Instruct, Evol-Instruct, Backtranslation and KG2Instruct workflows; its KG2Instruction example registers Qwen…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository exposes reusable code and examples for producing instruction-following data and prompt outputs with supported LLM backends, with Qwen-14B-Chat explicitly available in the model registry.

模型作用:Qwen Chat 14B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-14B-Chat acts as the chat LLM used by EasyInstruct workflows to generate or transform instruction/input/output examples and support KG-derived instruction creation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the project model registry plus README-described workflows; it is a concrete public toolkit, not a single customer story or benchmark-only result.

原始记录:EasyInstruct supports Qwen-14B-Chat for KG2Instruct and instruction-data generation workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

derisk-ai / OpenDerisk 使用 Qwen Chat 14B 处理研究分析和报告生成

derisk-ai / OpenDerisk · Qwen Chat 14B

A
厂商:Qwen / Alibaba 模型:Qwen Chat 14B 来源平台:github 最后复核:2026-06-27T07:30:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

derisk-ai / OpenDerisk 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:30:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenDerisk is a public AI-driven SRE and risk-intelligence framework; its core model configuration explicitly includes qwen-14b-chat mapped to the local Qwen-14B-Chat path and links to Qwen/Qwen-14B-Chat.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact provides an AI-native risk-intelligence/SRE framework, inherited DB-GPT-style model configuration, and local model slots that include Qwen-14B-Chat for assistant and analysis workflows.

模型作用:Qwen Chat 14B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen-14B-Chat serves as a supported local conversational model backend for risk-analysis, assistant and SRE workflows in the OpenDerisk framework.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from source configuration and project README rather than a named production deployment; still a public working framework with explicit Qwen-14B-Chat binding.

原始记录:OpenDerisk includes Qwen-14B-Chat as a local model for AI-driven SRE risk intelligence

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LibreChat / danny-avila 使用 Claude 2.0 处理软件工程任务执行

LibreChat / danny-avila · Claude 2.0

A
厂商:Anthropic / Claude 模型:Claude 2.0 来源平台:github 最后复核:2026-06-27T10:28:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LibreChat / danny-avila 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:28:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LibreChat is an open-source, self-hosted ChatGPT-style chat application with AI model selection across providers. Its public repository and product materials describe Anthropic/Claude support, and its runtime pricing/to…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact is the LibreChat GitHub repository and web product, an enhanced chat application with multi-provider model selection, presets, agents, message search, code interpreter and related chat features. The …

模型作用:Claude 2.0 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2 contributed the Anthropic large-language-model backend option for LibreChat conversations, supplying long-context natural-language reasoning and generation behind a self-hosted chat UI when users configured and…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the open-source product repository rather than a customer story. LibreChat is a multi-model application, so the public evidence confirms concrete Claude 2 support and integration metadata but does not p…

原始记录:LibreChat supported Claude 2 as an Anthropic model option for self-hosted multi-model chat

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

datawhalechina / Datawhale 使用 Qwen3 Max 处理研究分析和报告生成

datawhalechina / Datawhale · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:GitHub 最后复核:2026-06-27T07:30:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

datawhalechina / Datawhale 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:30:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:vibe-blog is a multi-agent AI writing assistant for deep research, outline planning, drafting, code integration, Mermaid diagrams, illustration, review, and formatting of long technical blog posts; its deployment exampl…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project README cites a public output artifact, Hello LLM-FineTuning, described as a 15+ chapter, 400k+ word technical tutorial with 100+ images produced using vibe-blog.

模型作用:Qwen3 Max 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 Max family is used as the text generation model behind the writing workflow, contributing planning, drafting, expansion, and polishing for the generated technical content.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public project README and configuration; the README uses qwen3-max-preview in the deployment example while the code whitelist includes qwen3-max, so this is bound to the Qwen3 Max family rather than…

原始记录:Datawhale vibe-blog uses Qwen3 Max family to generate long-form technical blog/tutorial content

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hikarioyama 使用 Step 3.7 Flash 处理软件工程任务执行

hikarioyama · Step 3.7 Flash

A
厂商:StepFun / Step 模型:Step 3.7 Flash 来源平台:github 最后复核:2026-06-27T07:37:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hikarioyama 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:37:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A standalone high-concurrency swarm harness uses Step-3.7-Flash as the model behind a conversational front door, planner, task DAG scheduler, persistent task queue, recall store, skills system, event log, and web UI.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a public CLI/TUI artifact named swarm that can launch demos, invaders specialist DAGs, or free-form planner goals while displaying live transcripts, server metrics, task events, and final results.

模型作用:Step 3.7 Flash 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Step-3.7-Flash is the core LLM used by the harness to route messages, plan goals, decompose them into worker lanes, generate replies, and drive multi-agent task execution.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub repository/README and code, but it is described by the maintainer as early work-in-progress rather than a production customer deployment.

原始记录:hikarioyama built a high-concurrency swarm-agent harness around Step-3.7-Flash

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Tunnello / Hello Agent team 使用 Qwen3 Max 处理软件工程任务执行

Tunnello / Hello Agent team · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:GitHub + Bilibili 最后复核:2026-06-27T07:30:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Tunnello / Hello Agent team 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:30:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ChatBI is a LangGraph-based conversational business-intelligence application that turns natural-language questions into SQLite SQL, executes queries, returns results, and generates visual charts through MCP tools; its A…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo includes screenshots and a Bilibili demo video for querying product categories, order trends, and chart visualizations from a local database through the ChatBI interface.

模型作用:Qwen3 Max 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 Max family acts as the agent LLM that routes tasks, reasons over schema retrieved by RAG, generates SQL, calls execution/chart tools, and composes the final streamed BI answer.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The exact selectable model shown in the API docs is qwen3-max-preview; this is treated as a Qwen3 Max family case. Demo video is public and reachable, but individual inference logs are not published.

原始记录:Tunnello ChatBI offers a Qwen3 Max Preview powered conversational BI agent for natural-language SQL and charts

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

sjusjdkzhsiqjqbz-rgb 使用 Step 3.7 Flash 处理软件工程任务执行

sjusjdkzhsiqjqbz-rgb · Step 3.7 Flash

A
厂商:StepFun / Step 模型:Step 3.7 Flash 来源平台:github 最后复核:2026-06-27T07:37:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

sjusjdkzhsiqjqbz-rgb 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:37:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Hermes Agent tool directly calls the StepFun chat completions API with video_url content blocks using the step-3.7-flash model, supporting local files, URLs, stepfile references, ffmpeg normalization, and reasoning_ef…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository ships stepfun_video.py and installation instructions so Hermes users can ask an agent to describe or summarize local/remote video clips and receive structured analysis responses.

模型作用:Step 3.7 Flash 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Step-3.7-Flash performs the multimodal video understanding and textual analysis after the tool prepares the video input, uploads or encodes it, and sends the prompt to StepFun's API.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public tool repository and README; examples are illustrative and the repository does not include independent production usage metrics.

原始记录:video-vision adds Step-3.7-Flash video analysis to Hermes Agent

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mino-park7 使用 Solar Open 100B 处理软件工程任务执行

mino-park7 · Solar Open 100B

A
厂商:Upstage / Solar 模型:Solar Open 100B 来源平台:github 最后复核:2026-06-27T07:38:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mino-park7 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Created a standalone vLLM plugin that packages Solar-specific changes for serving upstage/Solar-Open-100B with an unmodified upstream vLLM installation, including model registration, tool-call parsing, reasoning parsing…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Apache-2.0 GitHub repository and pip-installable package with serving commands and Python API examples for running upstage/Solar-Open-100B in vLLM.

模型作用:Solar Open 100B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Open 100B is the explicit target model; its architecture, reasoning blocks, chat-template token order, and parallel tool-call behavior drive the plugin components and serving workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public infrastructure artifact rather than a customer deployment; exact model binding is explicit in the README and serving examples.

原始记录:mino-park7 packaged Solar Open 100B support as an out-of-tree vLLM plugin

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Drenches 使用 Qwen3 Max 处理研究分析和报告生成

Drenches · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:GitHub 最后复核:2026-06-27T07:30:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Drenches 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:30:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:gov-doc-formatter is a Word/document formatting tool for Chinese party/government official documents. It uses a multi-agent workflow to judge writing style, clean non-official text, mark document structure, validate the…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public application accepts uploaded .doc/.docx files or pasted text and outputs a formatted Word document that users can download from the web interface or packaged Windows app.

模型作用:Qwen3 Max 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 Max supplies the LLM reasoning for routing, cleaning, structural marking, and validation before the formatter applies official-document styles to the output file.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public code and README document the workflow and qwen3-max default; no private user documents or run logs are exposed.

原始记录:Drenches gov-doc-formatter defaults to Qwen3 Max for official-document structure recognition and formatting

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

FAI-Solutions / Johannes Faber 使用 Step 3.7 Flash 处理软件工程任务执行

FAI-Solutions / Johannes Faber · Step 3.7 Flash

A
厂商:StepFun / Step 模型:Step 3.7 Flash 来源平台:github 最后复核:2026-06-27T07:37:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

FAI-Solutions / Johannes Faber 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:37:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project runs a local HTTP proxy between Continue in VS Code/VSCodium and NVIDIA NIM so Step 3.7 Flash can be used for chat, edit, apply, summarize, tool use, and image-input workflows without empty streaming respons…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides nim_proxy.py plus a concrete Continue config using model stepfun-ai/step-3.7-flash; the proxy strips incompatible request/stream fields and preserves tool_calls so Continue can display model outp…

模型作用:Step 3.7 Flash 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Step 3.7 Flash is the coding-assistant model served through NIM; the proxy adapts its streaming output format so Continue can consume its answers instead of silently discarding content.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public integration workaround repository; it documents a real compatibility issue and solution but not a named enterprise deployment.

原始记录:FAI-Solutions built a Continue/NIM proxy to make Step 3.7 Flash usable in VS Code

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

cyankiwi 使用 Solar Open 100B 处理真实任务执行

cyankiwi · Solar Open 100B

A
厂商:Upstage / Solar 模型:Solar Open 100B 来源平台:huggingface 最后复核:2026-06-27T07:38:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

cyankiwi 公开的真实任务执行案例,来源为 huggingface,复核于 2026-06-27T07:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Produced and published a 4-bit AWQ derivative of upstage/Solar-Open-100B for text-generation use, preserving the model card metadata that binds it to the Upstage base model.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A publicly accessible Hugging Face model artifact, cyankiwi/Solar-Open-100B-AWQ-4bit, with base_model metadata pointing to upstage/Solar-Open-100B and downloads recorded by Hugging Face.

模型作用:Solar Open 100B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Open 100B supplies the base weights and text-generation capability; the case adapts that model into a lower-bit deployment artifact.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Optimization artifact rather than an end-user product; exact base_model binding is explicit in Hugging Face metadata and README front matter.

原始记录:cyankiwi published a 4-bit AWQ Solar Open 100B variant for text generation

已有真实案例 真实任务执行huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

Liam-Frost 使用 Qwen3 Max 处理文档理解和结构化处理

Liam-Frost · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:GitHub 最后复核:2026-06-27T07:30:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Liam-Frost 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T07:30:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AutoApply is a local-first job-application automation workspace covering job discovery, fit scoring, tailored resume/cover-letter materials, form filling, human-gated submission, and application tracking. Its Qwen provi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact is a runnable Vue/FastAPI/PostgreSQL operator console that maintains job/application records, auditable agent traces, generated materials, and a human approval gate before submission.

模型作用:Qwen3 Max 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:When selected through the Qwen provider, Qwen3 Max contributes long-context reasoning over job descriptions and applicant memory, fit scoring explanations, tailored material generation, and workflow decisions inside the…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The repo proves product functionality and a pinned qwen3-max provider option; it does not publish a specific user submission trace, so evidence is strongest for supported real application use rather than a named employe…

原始记录:AutoApply exposes Qwen3 Max for local-first job discovery, fit scoring, and application-material generation

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

clawdbotatg 使用 Step 3.7 Flash 处理可玩交互原型构建

clawdbotatg · Step 3.7 Flash

A
厂商:StepFun / Step 模型:Step 3.7 Flash 来源平台:github 最后复核:2026-06-27T07:37:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

clawdbotatg 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T07:37:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A local arena used python-chess as referee and ran Step-3.7-Flash via llama.cpp Metal as White against DeepSeek-V4-Flash, loading one engine per move on a 128GB Mac because both models could not fit simultaneously.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository includes a public PGN and move log for a 2026-05-29 match where Step-3.7-Flash played legal moves as White and the recorded result was 1-0, White wins.

模型作用:Step 3.7 Flash 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:Step-3.7-Flash generated the chess moves for the White side from board-state prompts while python-chess enforced legality and recorded the game result.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an experimental public artifact rather than a production use case; it includes concrete outputs (PGN/log) and binds the exact model in the README and logs.

原始记录:clawd-local-ai-chess ran Step-3.7-Flash as a local chess-playing LLM

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AaryanK 使用 Solar Open 100B 处理软件工程任务执行

AaryanK · Solar Open 100B

A
厂商:Upstage / Solar 模型:Solar Open 100B 来源平台:huggingface 最后复核:2026-06-27T07:38:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

AaryanK 公开的代码代理与软件工程案例,来源为 huggingface,复核于 2026-06-27T07:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Converted Upstage Solar Open 100B into GGUF format and documented llama.cpp-oriented usage parameters for local or self-hosted text-generation workflows.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Hugging Face GGUF model repository with README description, base_model_relation quantized metadata, and files intended for GGUF-based inference stacks.

模型作用:Solar Open 100B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Open 100B is the source model being converted; the artifact exists to make its 102B MoE text-generation capability usable through GGUF runtimes.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Packaging and quantization case, not a downstream application; the Hugging Face model card explicitly names upstage/Solar-Open-100B as the source.

原始记录:AaryanK converted Solar Open 100B to GGUF for llama.cpp-style local inference

已有真实案例 代码代理与软件工程huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

SkyworkAI 使用 Qwen3 Max 处理软件工程任务执行

SkyworkAI · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:GitHub 最后复核:2026-06-27T07:30:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SkyworkAI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:30:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DeepResearchAgent is a hierarchical multi-agent runtime for deep-research and general task solving. Its LeetCode agent example supports command-line model selection and explicitly handles openrouter/qwen3-max as a non-v…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo provides runnable LeetCode agent scripts that parse programming tasks, generate code responses, and submit/evaluate answers via the included CodeSubmitter workflow.

模型作用:Qwen3 Max 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 Max can serve as the coding/reasoning model in the agent loop, producing reasoning and generated code for programming tasks while the framework manages task parsing, submission, and evaluation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence binds openrouter/qwen3-max in public example code and model manager; this is a framework example rather than a customer story, but it is a concrete runnable artifact rather than a benchmark-only leaderboard pag…

原始记录:SkyworkAI DeepResearchAgent includes Qwen3 Max as a selectable model for coding-agent LeetCode solving

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mlx-community 使用 Solar Open 100B 处理软件工程任务执行

mlx-community · Solar Open 100B

A
厂商:Upstage / Solar 模型:Solar Open 100B 来源平台:huggingface 最后复核:2026-06-27T07:38:41Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

mlx-community 公开的代码代理与软件工程案例,来源为 huggingface,复核于 2026-06-27T07:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Converted upstage/Solar-Open-100B into MLX format using mlx-lm 0.30.0 and published runnable load/generate instructions for Apple Silicon-oriented inference.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Hugging Face model repository, mlx-community/Solar-Open-100B-4bit, with MLX library metadata and example Python code for generating responses.

模型作用:Solar Open 100B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Open 100B provides the base model and generation behavior; the conversion makes that model available to MLX users in a 4-bit format.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Conversion artifact rather than production customer story; README and metadata explicitly bind the artifact to upstage/Solar-Open-100B.

原始记录:mlx-community converted Solar Open 100B to a 4-bit MLX model

已有真实案例 代码代理与软件工程huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

marksunner 使用 Step 3.7 Flash 处理软件工程任务执行

marksunner · Step 3.7 Flash

A
厂商:StepFun / Step 模型:Step 3.7 Flash 来源平台:github 最后复核:2026-06-27T07:37:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

marksunner 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:37:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository documents running StepFun Step 3.7 Flash locally on a single NVIDIA DGX Spark with llama.cpp, including use with a Hermes agent, file I/O tool calling, long-context settings, and stability tuning after su…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public notes report 26-27.5 tok/s generation, 90-107 tok/s prompt processing, a 96K token stable context ceiling, working tool calling, and a mitigation for CUDA graph crashes under sustained agent workloads.

模型作用:Step 3.7 Flash 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Step 3.7 Flash served as the local LLM backend for Hermes-style agent workflows, generating text, using tools, and handling multi-step research context on DGX Spark hardware.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a self-reported public engineering note with concrete measurements and configuration details; it is stronger as an engineering usage case than as a customer story.

原始记录:marksunner ran Step 3.7 Flash on a DGX Spark for Hermes agent use

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

massgen / MassGen 使用 Qwen3 Max 处理文档理解和结构化处理

massgen / MassGen · Qwen3 Max

A
厂商:Qwen / Alibaba 模型:Qwen3 Max 来源平台:GitHub 最后复核:2026-06-27T07:30:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

massgen / MassGen 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T07:30:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MassGen is an open-source multi-agent scaling system that coordinates multiple model-powered agents in the terminal to solve complex tasks through parallel attempts, observation, critique, iterative refinement, and voti…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact is the MassGen CLI/TUI, docs, and package that users can run for multi-agent evaluation, planning, specification writing, file-system operations, and general task solving.

模型作用:Qwen3 Max 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 Max can be one of the collaborating agents/backends, contributing independent solutions, critiques, refinements, and votes in the consensus workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public open-source integration and runnable system, not an external customer deployment; accepted because it names qwen3-max exactly and provides a concrete product artifact.

原始记录:MassGen supports Qwen3 Max as an Alibaba/Qwen backend for multi-agent collaborative task solving

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mike-ravkine 使用 Solar Open 100B 处理软件工程任务执行

mike-ravkine · Solar Open 100B

A
厂商:Upstage / Solar 模型:Solar Open 100B 来源平台:huggingface 最后复核:2026-06-27T07:38:41Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

mike-ravkine 公开的代码代理与软件工程案例,来源为 huggingface,复核于 2026-06-27T07:38:41Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Compressed upstage/Solar-Open-100B with llm-compressor 0.13.0 using a data-free FP8 recipe and documented a vLLM serving path with Solar-specific parsers and logits processors.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Hugging Face FP8 Dynamic model repository with README instructions for installing the Solar Open vLLM fork and serving the model on multi-GPU hardware.

模型作用:Solar Open 100B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Open 100B is the base model being compressed and served; its custom reasoning/tool-call handling is reflected in the required serving commands.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Compression/serving artifact, not an external product deployment; base_model metadata and README explicitly reference upstage/Solar-Open-100B.

原始记录:mike-ravkine compressed Solar Open 100B with llm-compressor for FP8 vLLM serving

已有真实案例 代码代理与软件工程huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

Yash2003Bisht 使用 DeepSeek-V2.5 处理软件工程任务执行

Yash2003Bisht · DeepSeek-V2.5

A
厂商:DeepSeek 模型:DeepSeek-V2.5 来源平台:github 最后复核:2026-06-27T07:39:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Yash2003Bisht 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:39:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository is a public VS Code extension for AI code completion. Its GitHub description states that 90% of the extension code was written by GPT, Claude and DeepSeek v2.5, binding DeepSeek-V2.5 to the concrete softw…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public VS Code extension repository named Null was published as the artifact for AI-powered code completion.

模型作用:DeepSeek-V2.5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 contributed code generation during implementation of the extension, alongside GPT and Claude.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public GitHub repository metadata/description rather than a detailed postmortem; no private deployment metrics are claimed.

原始记录:Null VS Code extension used DeepSeek V2.5 as part of AI-assisted code-completion development

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

otispm2020-dotcom 使用 DeepSeek-V2.5 处理软件工程任务执行

otispm2020-dotcom · DeepSeek-V2.5

A
厂商:DeepSeek 模型:DeepSeek-V2.5 来源平台:github 最后复核:2026-06-27T07:39:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

otispm2020-dotcom 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:39:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The public Chrome-extension repository describes HTML Slides Copilot as a tool that lets users control web browsers and extract information through natural-language commands by leveraging the SiliconFlow large-language-…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public browser-extension project was published for controlling HTML slide/web pages and extracting information from the browser through natural language.

模型作用:DeepSeek-V2.5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 is credited as the model used to translate human-readable instructions into executable JavaScript DOM operations.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The current README also documents newer/custom model providers; the DeepSeek-V2.5 binding is preserved in the repository's public GitHub description.

原始记录:HTML Slides Copilot used DeepSeek-V2.5 via SiliconFlow to turn natural-language commands into browser DOM operations

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Kitiara666 使用 DeepSeek-V2.5 处理软件工程任务执行

Kitiara666 · DeepSeek-V2.5

A
厂商:DeepSeek 模型:DeepSeek-V2.5 来源平台:github 最后复核:2026-06-27T07:39:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Kitiara666 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:39:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The public repository describes itself as a RunPod-optimized GPU client service for large models including DeepSeek V2.5, Qwen 3 and Kimi K2, with deployment files and service code for cloud GPU inference workflows.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Python/Shell/Docker repository was published containing a GPU client service, RunPod deployment files, example environment files and guides for running large models on cloud GPUs.

模型作用:DeepSeek-V2.5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-V2.5 is one of the target models the service is designed to support in the RunPod-optimized inference/deployment stack.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository metadata and file tree; no usage-volume or production-SLA claims are made.

原始记录:gpu-client-runpod packaged a RunPod-optimized GPU client service for DeepSeek V2.5 and other large models

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

gupta-v 使用 Gemini 1.5 Flash (May) 处理软件工程任务执行

gupta-v · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:github 最后复核:2026-06-27T07:43:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

gupta-v 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:43:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Streamlit application lets users upload videos and then chat with an AI agent about the video content, combining Gemini 1.5 Flash video understanding with DuckDuckGo web search through the Agno agent framework.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo provides a runnable app that processes uploaded video files and returns natural-language analysis and answers about the video content in an interactive chat interface.

模型作用:Gemini 1.5 Flash (May) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The application initializes Agno's Google Gemini model with id "gemini-1.5-flash" and uses it as the core multimodal model for analyzing uploaded video files and generating responses to user questions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public open-source application/repo rather than a named enterprise customer story; repository README and source code explicitly bind the artifact to Gemini 1.5 Flash.

原始记录:AI Video Analyzer & Chat Agent uses Gemini 1.5 Flash for uploaded-video understanding

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

jeevesh415 / nanobrowser 使用 Gemini 2.0 Flash Thinking exp. (Jan) 处理智能体流程编排

jeevesh415 / nanobrowser · Gemini 2.0 Flash Thinking exp. (Jan)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash Thinking exp. (Jan) 来源平台:GitHub 最后复核:2026-06-27T07:46:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

jeevesh415 / nanobrowser 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T07:46:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Nanobrowser, an open-source browser automation extension, added gemini-2.0-flash-thinking-exp-01-21 to its selectable models and upgraded LangChain Google GenAI dependencies so browser agents could use Gemini 2.0 Flash …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The PR states the model can be selected in settings, function calling works with Gemini 2.0 Flash Thinking, and model compatibility errors are reduced for Gemini users; the public artifact is the Nanobrowser extension l…

模型作用:Gemini 2.0 Flash Thinking exp. (Jan) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash Thinking exp. 01-21 was the reasoning model exposed to Nanobrowser users for agentic web automation steps that require tool/function calling.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub PR in a fork plus the public extension artifact; production usage volume is not disclosed.

原始记录:Nanobrowser enabled Gemini 2.0 Flash Thinking 01-21 for browser-agent function calling

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DEV-D-GR8 使用 Gemini 1.5 Flash (May) 处理知识检索和问答

DEV-D-GR8 · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:github 最后复核:2026-06-27T07:43:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DEV-D-GR8 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T07:43:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:TextGenius is a custom iOS keyboard extension and companion SwiftUI app that embeds a search/generation bar so users can ask questions or generate text for comments and replies from inside other iPhone apps.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact includes the iOS app source and a linked demo video; the README describes instant answers and generated comment/reply content delivered through the custom keyboard UI.

模型作用:Gemini 1.5 Flash (May) 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:The Swift GenerativeModelManager instantiates GenerativeModel with name "gemini-1.5-flash" and uses it to produce the keyboard's generated answers and text suggestions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from an open-source mobile app and demo rather than production usage metrics; source code explicitly names gemini-1.5-flash and the artifact is publicly reachable.

原始记录:TextGenius iOS keyboard uses Gemini 1.5 Flash for instant answers and reply generation

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

data5650 / Perplexica 使用 Gemini 2.0 Flash Thinking exp. (Jan) 处理代码审查和测试生成

data5650 / Perplexica · Gemini 2.0 Flash Thinking exp. (Jan)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash Thinking exp. (Jan) 来源平台:GitHub 最后复核:2026-06-27T07:46:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

data5650 / Perplexica 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T07:46:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Perplexica fork added gemini-2.0-flash-thinking-exp-01-21 to the Gemini provider configuration so users could select it in the UI for AI-powered search and answer generation.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The PR reports local Docker testing, confirms the model appears in the UI, is selectable, and works with sample queries.

模型作用:Gemini 2.0 Flash Thinking exp. (Jan) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash Thinking exp. 01-21 served as the answer-generation/reasoning model option for Perplexica search queries.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a fork PR and local testing notes rather than a merged upstream release; the repository artifact remains publicly reachable.

原始记录:Perplexica fork added Gemini 2.0 Flash Thinking 01-21 as a selectable search-answering model

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

shravya004 使用 Gemini 1.5 Flash (May) 处理软件工程任务执行

shravya004 · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:github 最后复核:2026-06-27T07:43:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

shravya004 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:43:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:PitchPerfectAI is a Streamlit web app where a job seeker enters role, experience, skills, achievements and tone preferences to generate a personalized cover letter and receive ATS match analysis.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo and README describe a working Streamlit workflow that produces polished cover letters plus an ATS score and improvement tips; the README also links a Google Drive demo video of the full workflow.

模型作用:Gemini 1.5 Flash (May) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The app initializes google.generativeai.GenerativeModel with "models/gemini-1.5-flash-latest" and uses Gemini 1.5 Flash as the LLM for cover-letter generation and analysis output.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source app/demo evidence, not an enterprise deployment; source code and README explicitly bind the project to Gemini 1.5 Flash and expose a reachable artifact.

原始记录:PitchPerfectAI uses Gemini 1.5 Flash to generate tailored cover letters and ATS feedback

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

arham-kk 使用 Gemini 2.0 Flash Thinking exp. (Jan) 处理软件工程任务执行

arham-kk · Gemini 2.0 Flash Thinking exp. (Jan)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash Thinking exp. (Jan) 来源平台:GitHub 最后复核:2026-06-27T07:46:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

arham-kk 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:46:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project implements an agent that decides whether a user request needs external context, retrieves relevant information with Google Search, and passes the gathered context plus the user request to Gemini 2.0 Flash Th…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes a working agent implementation and documents contextual reasoning, dynamic search integration, and asynchronous processing; the code sets thinking_model_id to gemini-2.0-flash-thinking-exp.

模型作用:Gemini 2.0 Flash Thinking exp. (Jan) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash Thinking provides the final reasoning and answer formulation over the retrieved search context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The code binds to the public Gemini 2.0 Flash Thinking experimental family without the 01-21 date suffix; accepted as the same Jan-era model family for this task.

原始记录:arham-kk built a search-augmented answering agent with Gemini 2.0 Flash Thinking

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Chanzhaoyu / chatgpt-web 使用 GPT-3.5 Turbo 处理真实任务执行

Chanzhaoyu / chatgpt-web · GPT-3.5 Turbo

A
厂商:OpenAI 模型:GPT-3.5 Turbo 来源平台:github 最后复核:2026-06-27T07:46:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Chanzhaoyu / chatgpt-web 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开源 ChatGPT Web 项目通过 OpenAI 官方 API 调用 gpt-3.5-turbo,为用户提供可自部署的网页聊天客户端,并支持 API Key、代理、Docker/Compose 等部署方式。

公开产物:产出一个可公开访问代码、可本地或容器化部署的 ChatGPT 网页应用;README 明确说明 ChatGPTAPI 通过 OpenAI official API 使用 gpt-3.5-turbo,且 OPENAI_API_MODEL 默认值为 gpt-3.5-turbo。

模型作用:GPT-3.5 Turbo 作为默认对话生成模型,负责理解用户消息并生成聊天回复,是该 Web 客户端核心交互能力的模型后端。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:开源自部署项目;是否生产部署取决于用户实例配置,但模型绑定、任务和公开代码产物均可核验。

原始记录:ChatGPT Web 用 GPT-3.5 Turbo 搭建开源 ChatGPT 网页客户端

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Yashwenth27 使用 Gemini 1.5 Flash (May) 处理软件工程任务执行

Yashwenth27 · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:github 最后复核:2026-06-27T07:43:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Yashwenth27 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:43:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:GitMe is a Streamlit developer tool that fetches files from a public GitHub repository, analyzes the codebase, and generates a polished markdown README for that repository.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo contains a runnable Streamlit app where users submit a GitHub profile/repo and receive generated README.md content covering overview, installation, usage, features and tech stack.

模型作用:Gemini 1.5 Flash (May) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The main application loads google.generativeai.GenerativeModel with model_name "models/gemini-1.5-flash" and calls it to understand project files and generate the README markdown.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an open-source developer tool rather than a customer story; the source code explicitly specifies Gemini 1.5 Flash and the repository is reachable.

原始记录:GitMe uses Gemini 1.5 Flash to generate README files from public GitHub repositories

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Abdulraqib20 使用 Gemini 2.0 Flash Thinking exp. (Jan) 处理软件工程任务执行

Abdulraqib20 · Gemini 2.0 Flash Thinking exp. (Jan)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash Thinking exp. (Jan) 来源平台:GitHub 最后复核:2026-06-27T07:46:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Abdulraqib20 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:46:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository describes an intelligent RAG system powered by Gemini 2.0 Flash Thinking, Qdrant vector storage, and Agno orchestration for uploading documents, processing web pages, rewriting queries, using web search f…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo provides installation/configuration instructions and a concrete RAG application artifact for multi-PDF and web-content Q&A.

模型作用:Gemini 2.0 Flash Thinking exp. (Jan) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash Thinking is the LLM reasoning layer used to rewrite queries and generate AI-assisted answers from retrieved document/web context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:README names Gemini 2.0 Flash Thinking rather than the exact 01-21 suffix; no deployment metrics are disclosed.

原始记录:Abdulraqib20 built an Agentic RAG app powered by Gemini 2.0 Flash Thinking

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

talkingwallace / ChatGPT-Paper-Read… 使用 GPT-3.5 Turbo 处理研究分析和报告生成

talkingwallace / ChatGPT-Paper-Reader · GPT-3.5 Turbo

A
厂商:OpenAI 模型:GPT-3.5 Turbo 来源平台:github 最后复核:2026-06-27T07:46:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

talkingwallace / ChatGPT-Paper-Reader 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:46:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开源工具将 PDF 学术论文分段,调用 gpt-3.5-turbo 对每一部分生成摘要,并基于全部分段摘要回答用户针对论文的问题。

公开产物:产出论文分段摘要、针对作者/方法/性能指标/数据集等问题的结构化回答,以及可交互的论文问答接口;README 明确写明该仓库使用 gpt-3.5-turbo 本地阅读 PDF 论文。

模型作用:GPT-3.5 Turbo 负责对论文文本片段进行理解、压缩摘要、跨片段衔接上下文,并生成最终问答结果。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:学术辅助阅读场景,结果需由用户核对原论文;公开仓库和 README 对模型、任务、输出流程均有直接说明。

原始记录:ChatGPT-Paper-Reader 用 GPT-3.5 Turbo 阅读和总结 PDF 学术论文

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hyunwen 使用 Gemini 2.0 Flash Thinking exp. (Jan) 处理软件工程任务执行

hyunwen · Gemini 2.0 Flash Thinking exp. (Jan)

A
厂商:Google / Gemini 模型:Gemini 2.0 Flash Thinking exp. (Jan) 来源平台:GitHub 最后复核:2026-06-27T07:46:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hyunwen 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:46:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project uses Gemini 2.0 Flash Thinking as a teacher model for C/C++ vulnerability detection, generating step-by-step reasoning and vulnerability assessments for code snippets that are then used to train smaller stud…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes scripts and documentation for producing JSON records containing code-snippet identifiers, model thinking/reasoning, final vulnerability responses, and labels for downstream distillation.

模型作用:Gemini 2.0 Flash Thinking exp. (Jan) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.0 Flash Thinking supplies the reasoning traces and final vulnerability judgments that form the teacher signal for knowledge distillation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The code uses gemini-2.0-flash-thinking-exp without a date suffix; the repository is an applied research artifact rather than a deployed commercial product.

原始记录:hyunwen used Gemini 2.0 Flash Thinking to generate reasoning data for C/C++ vulnerability detection distillation

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AdamXacur 使用 Gemini 1.5 Flash (May) 处理代码审查和测试生成

AdamXacur · Gemini 1.5 Flash (May)

A
厂商:Google / Gemini 模型:Gemini 1.5 Flash (May) 来源平台:github 最后复核:2026-06-27T07:43:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AdamXacur 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T07:43:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:CallCenterQualityEvaluator.py is a desktop application for call-center QA teams that takes audio/transcription inputs, applies configurable quality rubrics, and evaluates agent performance for a named company.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public application generates dialization text, evaluates call-center agent performance against quality rubrics, adds a score at the end of the evaluation, tracks token usage, and includes a linked YouTube usage tuto…

模型作用:Gemini 1.5 Flash (May) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:The Python application calls google.generativeai.GenerativeModel with model_name "gemini-1.5-flash" in its text generation, evaluation and polishing functions; Gemini 1.5 Flash produces the agent-quality evaluation text…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public open-source desktop app rather than an official customer story; source code explicitly names gemini-1.5-flash and the repository plus tutorial artifact are reachable.

原始记录:CallCenterQualityEvaluator uses Gemini 1.5 Flash to score call-center agent performance from transcripts

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yuezk / chatgpt-mirror 使用 GPT-3.5 Turbo 处理真实任务执行

yuezk / chatgpt-mirror · GPT-3.5 Turbo

A
厂商:OpenAI 模型:GPT-3.5 Turbo 来源平台:github 最后复核:2026-06-27T07:46:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yuezk / chatgpt-mirror 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ChatGPT Mirror 是一个可自部署的 ChatGPT 镜像网页应用,README 明确标注基于 gpt-3.5-turbo,并要求用户配置 OpenAI API Key 后运行服务。

公开产物:产出一个 Node.js 自托管聊天 Web 应用,用户在本地或服务器启动后可通过浏览器访问并与 GPT-3.5 Turbo 对话。

模型作用:GPT-3.5 Turbo 是项目声明的基础模型,承担聊天问答中的自然语言理解和回复生成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:仓库已归档,但公开代码和 README 仍可访问;属于真实开源应用产物,不是教程或集合页。

原始记录:ChatGPT Mirror 基于 GPT-3.5 Turbo 提供自托管聊天镜像

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

eosphoros-ai / DB-GPT 使用 Qwen1.5 Chat 110B 处理代码审查和测试生成

eosphoros-ai / DB-GPT · Qwen1.5 Chat 110B

A
厂商:Qwen / Alibaba 模型:Qwen1.5 Chat 110B 来源平台:GitHub 最后复核:2026-06-27T07:46:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

eosphoros-ai / DB-GPT 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DB-GPT maintainers added and tested Qwen1.5-110B-Chat support in the DB-GPT framework, with the PR explicitly testing LLM_MODEL=qwen1.5-110b-chat and attaching a working snapshot.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The PR added Qwen110B support to DB-GPT and documented testing with LLM_MODEL=qwen1.5-110b-chat, making the model selectable as a backend for DB-GPT workflows.

模型作用:Qwen1.5 Chat 110B 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen1.5-110B-Chat served as the large chat model backend DB-GPT integrated for conversational/database-oriented LLM application workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a product integration PR rather than an end-customer story; still includes concrete model identifier, maintainer organization, tested configuration, and public artifact.

原始记录:DB-GPT integrates Qwen1.5-110B-Chat as a supported LLM backend

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

binary-husky / gpt_academic 使用 GPT-3.5 Turbo 处理研究分析和报告生成

binary-husky / gpt_academic · GPT-3.5 Turbo

A
厂商:OpenAI 模型:GPT-3.5 Turbo 来源平台:github 最后复核:2026-06-27T07:46:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

binary-husky / gpt_academic 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:46:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:GPT 学术优化为学术场景提供论文润色纠错、公式显示、项目代码解析、多模型混合调用等功能,README 明确包含 OpenAI GPT-3.5,并在可用模型配置中列出 gpt-3.5-turbo。

公开产物:产出可运行的学术辅助 Web/GUI 工具,用于论文阅读润色、问答、公式处理和代码项目分析;README 展示了论文/代码处理输出截图和 GPT-3.5 混合调用能力。

模型作用:GPT-3.5 Turbo 作为可用 OpenAI 模型之一,为润色、问答、摘要和代码解释等文本生成任务提供核心语言理解与生成能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目支持多模型,单次运行可由用户配置模型;README 和配置文件共同证明 gpt-3.5-turbo 是公开支持的真实模型选项。

原始记录:GPT 学术优化用 GPT-3.5 Turbo 支持论文阅读、润色和代码解析

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

InternLM/lmdeploy GitHub user commu… 使用 Qwen1.5 Chat 110B 处理真实任务执行

InternLM/lmdeploy GitHub user community · Qwen1.5 Chat 110B

A
厂商:Qwen / Alibaba 模型:Qwen1.5 Chat 110B 来源平台:GitHub 最后复核:2026-06-27T07:46:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

InternLM/lmdeploy GitHub user community 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A user deployed PATH/Qwen1.5-110B-Chat with lmdeploy serve api_server, model name qwen, tensor parallelism 8 or 4, then called /v1/chat/completions with a chat prompt.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The issue records a concrete deployment attempt on four A800-SXM4-80GB GPUs: Qwen1.5-110B-Chat service started, but the chat completion request did not return promptly while the 72B service did.

模型作用:Qwen1.5 Chat 110B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen1.5-110B-Chat was the hosted chat model behind the OpenAI-compatible API endpoint for answering user prompts.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a troubleshooting report with a negative serving outcome, not a polished showcase; it is still a real public deployment/use artifact with exact command, hardware, and model path.

原始记录:LMDeploy user serves Qwen1.5-110B-Chat as an OpenAI-compatible chat API on A800 GPUs

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

harry0703 / MoneyPrinterTurbo 使用 GPT-3.5 Turbo 处理多模态内容处理

harry0703 / MoneyPrinterTurbo · GPT-3.5 Turbo

A
厂商:OpenAI 模型:GPT-3.5 Turbo 来源平台:github 最后复核:2026-06-27T07:46:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

harry0703 / MoneyPrinterTurbo 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T07:46:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:MoneyPrinterTurbo 是一键生成高清短视频的开源项目,功能列表包括 AI 自动生成视频文案、批量视频生成、字幕、配音、背景音乐和素材检索;配置示例中公开列出 gpt-3.5-turbo/gpt-35-turbo 模型名称。

公开产物:产出可播放的竖屏/横屏短视频 Demo 和可部署的 Web/API 视频生成系统,支持从主题生成文案、合成语音、匹配素材、生成字幕并导出视频。

模型作用:GPT-3.5 Turbo 在该流程中作为文本生成模型选项,用于把用户主题或需求扩展为视频脚本/文案,驱动后续配音、素材检索和视频合成步骤。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目支持多家模型提供方,具体运行模型由配置决定;公开配置示例直接列出 GPT-3.5 Turbo,README 证明其真实短视频生成产物和任务流程。

原始记录:MoneyPrinterTurbo 用 GPT-3.5 Turbo 生成短视频文案并自动产出高清视频

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

vllm-project/vllm GitHub user commu… 使用 Qwen1.5 Chat 110B 处理软件工程任务执行

vllm-project/vllm GitHub user community · Qwen1.5 Chat 110B

A
厂商:Qwen / Alibaba 模型:Qwen1.5 Chat 110B 来源平台:GitHub 最后复核:2026-06-27T07:46:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

vllm-project/vllm GitHub user community 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A user used vLLM 0.4.2 AsyncLLMEngine to run Qwen1.5-110B-Chat-GPTQ-Int4 from ModelScope on A100 80G, tuning gpu_memory_utilization, max_model_len, max_num_seqs, and sending 330-token prompts with short outputs.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The report gives observed latency/QPS figures: about 0.45s per request at QPS=1, 0.70s at QPS=2, 1.4s at QPS=4, and 2.9s at QPS=8, with effective batch size around 2.

模型作用:Qwen1.5 Chat 110B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen1.5-110B-Chat-GPTQ-Int4 was the inference model being served by vLLM for low-latency short-answer generation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a performance/usage issue rather than a customer case study; it includes exact model URL, serving framework, hardware, parameters, and measured outputs.

原始记录:vLLM user runs Qwen1.5-110B-Chat-GPTQ-Int4 inference on A100 and measures request latency

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hiyouga/LLaMA-Factory GitHub user c… 使用 Qwen1.5 Chat 110B 处理真实任务执行

hiyouga/LLaMA-Factory GitHub user community · Qwen1.5 Chat 110B

A
厂商:Qwen / Alibaba 模型:Qwen1.5 Chat 110B 来源平台:GitHub 最后复核:2026-06-27T07:46:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hiyouga/LLaMA-Factory GitHub user community 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A user configured LLaMA-Factory SFT with DeepSpeed ZeRO-3 offload to fine-tune ../../model/qwen/Qwen1.5-110B-Chat on a custom dataset named Kee_Instruction_NewEstabalish, using LoRA rank 128, cutoff_len 6000, and 4 A40 …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public issue includes the training command, dataset/configuration, expected goal of using 4×A40 plus DeepSpeed ZeRO-3 to fine-tune Qwen1.5-110B, and the encountered device placement runtime error.

模型作用:Qwen1.5 Chat 110B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen1.5-110B-Chat was the base instruction/chat model being adapted through SFT/LoRA for the user's domain dataset.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a fine-tuning troubleshooting artifact, not a successful public demo; however it documents a concrete organization/user workflow with exact model path, task, dataset, and training setup.

原始记录:LLaMA-Factory user attempts SFT fine-tuning of Qwen1.5-110B-Chat with DeepSpeed ZeRO-3 on 4×A40

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xorbitsai/inference GitHub user com… 使用 Qwen1.5 Chat 110B 处理真实任务执行

xorbitsai/inference GitHub user community · Qwen1.5 Chat 110B

A
厂商:Qwen / Alibaba 模型:Qwen1.5 Chat 110B 来源平台:GitHub 最后复核:2026-06-27T07:46:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xorbitsai/inference GitHub user community 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A user launched Qwen1.5-110B-Chat-GPTQ-Int4 through Xinference with model_uid qwen1.5-110b-4090, model_format gptq, quantization Int4, n_gpu=4, replica=2, gpu_memory_utilization=0.9, max_model_len=12000, and model_size_…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The issue records the serving logs, including caching qwen/Qwen1.5-110B-Chat-GPTQ-Int4 from ModelScope and loading the model with vLLM configuration before an NCCL error when using replicas.

模型作用:Qwen1.5 Chat 110B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen1.5-110B-Chat-GPTQ-Int4 was the large chat model being exposed through Xinference/vLLM for local multi-GPU serving.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an operational bug report and not a success story; it remains a real public deployment artifact with exact model identifier, serving parameters, hardware, and logs.

原始记录:Xinference user launches replicated Qwen1.5-110B-Chat-GPTQ-Int4 serving on 8×4090 GPUs

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

arunjayzui 使用 MiniMax M1 80k 处理真实任务执行

arunjayzui · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:Hugging Face Spaces 最后复核:2026-06-27T07:52:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

arunjayzui 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T07:52:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上用 Gradio 搭建 MiniMaxAI/MiniMax-M1-80k 的交互式文本生成/聊天推理 Demo,并通过 Hugging Face 登录调用 novita inference provider。

公开产物:公开可访问的 Hugging Face Space;app.py 明确写有 gr.load("models/MiniMaxAI/MiniMax-M1-80k", accept_token=button, provider="novita"),用于向用户展示该模型的在线推理能力。

模型作用:MiniMax M1 80k 是 Space 载入并对外提供文本生成/聊天响应的核心模型,负责根据用户输入生成回复。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开 Demo/推理 Space,不是生产客户故事;但代码和 Space 元数据均绑定 exact model MiniMaxAI/MiniMax-M1-80k,artifact 可访问。

原始记录:arunjayzui 发布 MiniMax-M1-80k Hugging Face Gradio 推理 Demo

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

yYorky 使用 Gemini 2.5 Pro (Mar) 处理研究分析和报告生成

yYorky · Gemini 2.5 Pro (Mar)

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro (Mar) 来源平台:github 最后复核:2026-06-27T07:53:33Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yYorky 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:53:33Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Analyze football match highlight videos from YouTube, identify teams and players, rate player performances, summarize key moments and attach timestamped tactical observations.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A Streamlit football analysis tool with demo GIFs and source code; the README states it is powered by Gemini 2.5 Pro and specifically notes use of gemini-2.5-pro-exp-03-25 in modules/config.py.

模型作用:Gemini 2.5 Pro (Mar) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro processes the YouTube video/highlight input and generates detailed match insights, player ratings, key moments and timestamp-linked analysis.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public repo is reachable and exact March experimental model string is named in README; evidence is project-maintainer supplied.

原始记录:FootballVideoAnalyst uses Gemini 2.5 Pro Experimental to analyze YouTube football highlights

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

JPaul02 使用 MiniMax M1 80k 处理真实任务执行

JPaul02 · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:Hugging Face Spaces 最后复核:2026-06-27T07:52:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

JPaul02 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T07:52:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上发布一个面向登录用户的 MiniMaxAI/MiniMax-M1-80k Gradio 推理页面,用于体验该模型的文本生成/聊天能力。

公开产物:公开 Space 页面可访问;app.py 的侧栏文案写明 This Space showcases the MiniMaxAI/MiniMax-M1-80k model,并通过 gr.load 载入 models/MiniMaxAI/MiniMax-M1-80k。

模型作用:MiniMax M1 80k 作为 Space 的唯一载入模型,承担用户提示词到文本回复的生成任务。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开 Demo/推理 Space,未证明生产落地;但使用者、任务、产物 URL 和 exact model 绑定清晰。

原始记录:JPaul02 发布 MiniMax-M1-80k Hugging Face Gradio 推理 Demo

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

Ekaterina Ya / lastminute_legal 使用 Gemini 2.5 Pro (Mar) 处理研究分析和报告生成

Ekaterina Ya / lastminute_legal · Gemini 2.5 Pro (Mar)

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro (Mar) 来源平台:github 最后复核:2026-06-27T07:53:33Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Ekaterina Ya / lastminute_legal 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T07:53:33Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Operate a Telegram bot that analyzes text and image advertising creatives against Russian advertising legislation, using RAG over Federal Antimonopoly Service case data to provide compliance feedback.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Telegram-bot codebase for advertising-law compliance checks; the README states it uses Gemini 2.5 Pro API for legal reasoning and combines Gemini 2.5 Pro with RAG for legal analysis.

模型作用:Gemini 2.5 Pro (Mar) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro supplies the legal-reasoning generation step after retrieval, producing the final compliance analysis and explanations from the creative and relevant FAS cases.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository README provides the primary evidence; legal advice quality is not independently validated, but the artifact and task are concrete and public.

原始记录:lastminute_legal uses Gemini 2.5 Pro API plus RAG to check Russian advertising creatives for compliance

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Udayxyz 使用 MiniMax M1 80k 处理真实任务执行

Udayxyz · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:Hugging Face Spaces 最后复核:2026-06-27T07:52:51Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Udayxyz 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T07:52:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上搭建标题/描述为 Text generator 的 MiniMaxAI/MiniMax-M1-80k Gradio 应用,供用户登录后调用模型做文本生成。

公开产物:公开 Space 页面可访问;README 元数据标注 short_description 为 Text generator,app.py 明确通过 novita provider 载入 models/MiniMaxAI/MiniMax-M1-80k。

模型作用:MiniMax M1 80k 是该文本生成 Space 的核心推理模型,负责根据用户输入生成输出文本。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开文本生成 Demo/Space,非客户生产案例;但 exact model、使用者、任务和公开 artifact 均可核验。

原始记录:Udayxyz 发布 MiniMax-M1-80k Text generator Space

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

Jack Perry / Ocean County Hive Mind 使用 Gemini 2.5 Pro (Mar) 处理软件工程任务执行

Jack Perry / Ocean County Hive Mind · Gemini 2.5 Pro (Mar)

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro (Mar) 来源平台:github 最后复核:2026-06-27T07:53:33Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Jack Perry / Ocean County Hive Mind 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T07:53:33Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Next.js/React website to automate creating thumbnails for Ocean County Hive Mind Magic: The Gathering content.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public ochm-thumbnails repository containing the thumbnail-automation website source; the README states it was created with Cursor and gemini-2.5-pro-exp-03-25 and links to the Ocean County Hive Mind YouTube channel w…

模型作用:Gemini 2.5 Pro (Mar) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro Experimental was used through Cursor as the coding model to help implement the thumbnail-generation website and automate the production workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence explicitly names the March experimental model string; output usage is inferred from the repository README's link to the public YouTube channel, while the code artifact itself is public and reachable.

原始记录:Ocean County Hive Mind thumbnail automation website was created with Cursor and gemini-2.5-pro-exp-03-25

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

emilykangdev / stower 使用 Qwen3 Max Thinking 处理代码审查和测试生成

emilykangdev / stower · Qwen3 Max Thinking

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking 来源平台:github 最后复核:2026-06-27T08:05:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

emilykangdev / stower 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:05:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Stower 在 ship-with-fusion 门禁中把 Qwen3 Max Thinking 作为第四谱系的独立 judge,对 Kimi、Grok、DeepSeek 等非同源 panel 产生的代码审计发现做二阶段复核,判断哪些问题真实成立、哪些属于噪声或幻觉。

公开产物:PR 记录显示 Qwen3 Max Thinking judge 在第 1 轮确认 3 个 P2 问题,其中 2 个已修复、1 个转入 TODO;第 2 轮继续验证修复成立,并判定后续 1 个 P1 与 4 个 P2 阻塞项均未通过源代码核验,避免把 hallucinated finding 当成阻塞问题。

模型作用:Qwen3 Max Thinking 的贡献是作为独立推理裁判读取约 124k tokens 的分支上下文、plans、docs、源码和测试结果,交叉核验 panel 发现并输出可执行的 confirmed/noise 判定。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub PR 描述;未看到完整私有运行日志,但 PR 明确列出 judge 模型、任务、轮次和审计结果。

原始记录:Stower 用 Qwen3 Max Thinking 作为独立 Judge 复核多模型代码审计结论

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

IrisY2012 使用 Qwen3 235B A22B 2507 处理软件工程任务执行

IrisY2012 · Qwen3 235B A22B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B A22B 2507 来源平台:github 最后复核:2026-06-27T08:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

IrisY2012 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:创建一个 C++ 命令行密码生成器,根据输入字符串用 FNV-1a 64 位哈希生成确定性随机种子,再生成包含大小写字母、数字和特殊字符的 12 位密码。GitHub 仓库描述明确标注 Made with AI(Qwen3-235B-A22B-2507)。

公开产物:公开仓库包含可核验的 passwordCreate.cpp 源码,实现 PasswordCreater::create、字符集约束、随机打乱和 main 输入输出流程。

模型作用:Qwen3-235B-A22B-2507 被仓库作者标注为 AI 生成来源,贡献了核心 C++ 代码结构、哈希种子方案和密码生成逻辑。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人开源仓库,证据来自 GitHub 仓库标题/描述和公开源码;未显示生产部署。

原始记录:IrisY2012 用 Qwen3-235B-A22B-2507 生成确定性密码生成器

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Fireworks AI 使用 Qwen2 72B 处理智能体流程编排

Fireworks AI · Qwen2 72B

A
厂商:Qwen / Alibaba 模型:Qwen2 72B 来源平台:official_web 最后复核:2026-06-27T08:07:03Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Fireworks AI 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T08:07:03Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Fireworks AI productized Qwen2 72B Instruct as an API and interactive playground for instruction-following text generation, natural-language understanding and coding or reasoning style prompts.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Fireworks model page exposes Qwen2 72B Instruct as a serverless model endpoint/playground that developers can use for text-generation workloads.

模型作用:Qwen2 72B 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2 72B Instruct is the underlying large language model that generates the assistant responses served through the Fireworks endpoint.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public provider product page rather than a named end-customer story; page explicitly identifies Qwen2 72B Instruct and the API/playground artifact.

原始记录:Fireworks AI provides Qwen2 72B Instruct as a hosted API and playground

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

Siavarshan21 使用 Qwen3 235B A22B 2507 处理浏览器 3D 世界构建

Siavarshan21 · Qwen3 235B A22B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B A22B 2507 来源平台:github 最后复核:2026-06-27T08:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Siavarshan21 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-27T08:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:为餐厅或咖啡馆制作现代响应式落地页,展示菜单、环境、图库和预订入口。GitHub 仓库描述含 profile Qwen3-235B-A22B-2507,并保留模型生成的项目描述。

公开产物:公开仓库包含 React/TypeScript/Vite 应用源码,client/pages/Index.tsx 实现 Saveur Exquisite 餐厅主页、3D 菜品展示、导航、菜单/图库/预订等页面区块。

模型作用:Qwen3-235B-A22B-2507 参与生成了项目描述和落地页实现,包括页面文案、组件结构和前端交互布局。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人 GitHub 项目,证据来自仓库描述与可访问源码;未发现独立线上部署链接。

原始记录:Siavarshan21 用 Qwen3-235B-A22B-2507 构建餐厅 Landing Page

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MinhHaDuong / aedist-technical-repo… 使用 Qwen3 Max Thinking 处理软件工程任务执行

MinhHaDuong / aedist-technical-report · Qwen3 Max Thinking

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking 来源平台:github 最后复核:2026-06-27T08:05:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MinhHaDuong / aedist-technical-report 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:05:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:AEDIST technical report 项目在 PR #316 中新增脚本和审计产物,使用包含 Qwen3 Max Thinking 在内的多模型 panel 审计 docs/argument.md,并保存每个模型的原始 markdown/json 响应,用于后续假设提取和论证修订。

公开产物:公开产物 docs/audit-responses/qwen_qwen3-max-thinking.md 显示该模型在 2026-05-01 运行,处理 4356 input tokens / 327 output tokens,用约 10.8 秒输出最强不一致、缺失桥接证据和建议的最小修正等审计内容。

模型作用:Qwen3 Max Thinking 贡献了对研究论证的结构化批判性推理,指出 deep-research cell 与 F1 饱和论断之间的不一致、缺少可验证桥接证据,并给出最小修正建议,成为 PR 中提交的可复核审计响应之一。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据包含公开 PR 和可访问的原始模型响应文件;这是研究论证审计产物,不是 benchmark 或教程。

原始记录:AEDIST technical report 用 Qwen3 Max Thinking 产出论文论证审计原始报告

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenRouter 使用 Qwen2 72B 处理真实任务执行

OpenRouter · Qwen2 72B

A
厂商:Qwen / Alibaba 模型:Qwen2 72B 来源平台:official_web 最后复核:2026-06-27T08:07:03Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

OpenRouter 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T08:07:03Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenRouter made Qwen 2 72B Instruct available through its model routing API so applications can call the model for multilingual language understanding, coding, mathematics and reasoning prompts.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public OpenRouter model page provides the Qwen 2 72B Instruct endpoint with API/pricing metadata and links back to the Qwen Hugging Face model.

模型作用:Qwen2 72B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2 72B supplies the text-generation and reasoning capability that OpenRouter routes to developers through its unified API.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public model endpoint/product page, not an individual customer story; exact model identity is bound by the OpenRouter slug and Hugging Face link on the page.

原始记录:OpenRouter lists Qwen 2 72B Instruct for routed API access

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

Muneeb007j 使用 Qwen3 235B A22B 2507 处理真实任务执行

Muneeb007j · Qwen3 235B A22B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B A22B 2507 来源平台:github 最后复核:2026-06-27T08:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Muneeb007j 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:围绕原创 Div1/Div2 级别竞赛题 Prime Factorial Array,使用 Qwen3-235B-A22B-2507 在关闭 thinking 的条件下进行三次求解尝试,并记录模型输出思路与失败测试。

公开产物:公开仓库包含题面、测试用例、参考解法和 qwen/conversations.md;记录显示两条公开 Qwen Chat 分享链接及三次尝试,其中尝试 1 失败测试 1/2/3/5,尝试 2 失败测试 2/3/5,尝试 3 失败全部 5 个测试。

模型作用:Qwen3-235B-A22B-2507 生成了三版解题方案/代码尝试,仓库作者将其作为对抗评测对象,用于验证题目能否诱导大模型给出错误贪心解。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是公开项目中的对抗性使用案例,结果为失败样例而非成功交付;仍具备具体模型、使用者、任务、公开产物和原始对话证据。

原始记录:Muneeb007j 用 Qwen3-235B-A22B-2507 进行竞赛题对抗测试

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chronosllc0-ai / aegis-ui-agent 使用 Qwen3 Max Thinking 处理智能体流程编排

chronosllc0-ai / aegis-ui-agent · Qwen3 Max Thinking

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking 来源平台:github 最后复核:2026-06-27T08:05:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chronosllc0-ai / aegis-ui-agent 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:05:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Aegis UI Agent 是一个通过自然语言理解并操作任意网站的通用 UI navigator;PR #79 为 OpenRouter 等 provider 接入 reasoning_delta 流,并在模型选择器中把 Qwen3 Max Thinking 标为支持 reasoning 的模型。

公开产物:PR 描述列出后端 provider 流式 reasoning_delta、WebSocket reasoning_delta 事件、/reason on|off|low|medium|high|status 命令、前端 ThinkingCard、PlusMenu reasoning toggle、ActionLog reasoning dropdown,以及包含 Qwen3 Max Thinking 的 reasoning-capable model picker。

模型作用:Qwen3 Max Thinking 被绑定为该代理产品可选的长推理模型,贡献在于为复杂网页导航/操作任务提供可展示的思考过程和 reasoning effort 控制,使用户能在 UI 中查看与折叠模型推理流。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据是公开产品仓库 PR;PR 明确列出 Qwen3 Max Thinking 的 reasoning 能力接入,但未附单次终端运行日志。

原始记录:Aegis UI Agent 在通用网页操作代理中为 Qwen3 Max Thinking 接入 reasoning/thinking 流式展示

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

akhaliq 使用 Qwen3 235B A22B 2507 处理真实任务执行

akhaliq · Qwen3 235B A22B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B A22B 2507 来源平台:huggingface_spaces 最后复核:2026-06-27T08:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

akhaliq 公开的真实任务执行案例,来源为 huggingface_spaces,复核于 2026-06-27T08:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上创建 Gradio ChatInterface,让用户通过 Hugging Face InferenceClient 和 fireworks-ai provider 与 Qwen/Qwen3-235B-A22B-Instruct-2507 进行多轮聊天。

公开产物:公开 Space 配置 sdk=gradio、app_file=app.py;源码明确设置 MODEL = "Qwen/Qwen3-235B-A22B-Instruct-2507",并流式返回 chat.completions.create 的回答。

模型作用:Qwen3-235B-A22B-Instruct-2507 是该 Space 的核心生成模型,负责根据用户输入和对话历史生成聊天回复。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:社区演示应用,属于公开可访问 demo;不是客户生产案例。

原始记录:akhaliq 在 Hugging Face Space 搭建 Qwen3-235B-A22B-Instruct-2507 Gradio 聊天 Demo

已有真实案例 真实任务执行huggingface_spacesA 类可核验real_case auto_approved 进入模型卡精选

NVIDIA 使用 Qwen2 72B 处理真实任务执行

NVIDIA · Qwen2 72B

A
厂商:Qwen / Alibaba 模型:Qwen2 72B 来源平台:official_web 最后复核:2026-06-27T08:07:03Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

NVIDIA 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T08:07:03Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:NVIDIA Build exposes Qwen2 72B Instruct as a Qwen-published model page for developers to test and deploy text-generation inference in NVIDIA's model catalog environment.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public NVIDIA Build artifact exists under the qwen/qwen2-72b-instruct route, presenting the model for API-style usage and deployment exploration.

模型作用:Qwen2 72B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2 72B Instruct is the language model performing the instruction-following generation behind the NVIDIA Build model entry.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public model-catalog/product artifact rather than a customer case study; exact model is bound by the NVIDIA Build URL path qwen/qwen2-72b-instruct.

原始记录:NVIDIA Build publishes Qwen2 72B Instruct as a deployable model endpoint

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

2witstudios / PageSpace 使用 Qwen3 Max Thinking 处理智能体流程编排

2witstudios / PageSpace · Qwen3 Max Thinking

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking 来源平台:github 最后复核:2026-06-27T08:05:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

2witstudios / PageSpace 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:05:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:PageSpace 在 PR #747 中为其 AI chat / agent integration 系统增加多家最新模型支持,其中 Qwen(OpenRouter) 条目明确包含 Qwen3 Max Thinking,并要求验证 AI chat 的 model selectors 中新模型可见。

公开产物:PR 总结显示新增 Anthropic、Google、OpenAI、OpenRouter/Qwen 等模型,并把 Qwen3 Max Thinking 纳入 Qwen OpenRouter 支持列表;验证步骤要求检查 AI chat 模型选择器中新模型出现,同时配套 Slack/Notion adapters、agent integration 面板和审计日志等产品功能改动。

模型作用:Qwen3 Max Thinking 的贡献是作为 PageSpace AI Chat 中可选的 Qwen/OpenRouter 高推理模型,为用户在工作流和集成上下文中执行复杂问答、任务协助和 agent 操作提供模型后端。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub PR;这是产品模型接入与可用性验证,不是独立 customer story。

原始记录:PageSpace 将 Qwen3 Max Thinking 接入 AI Chat 模型选择器和 OpenRouter 路由

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lucataco 使用 Qwen2 72B 处理软件工程任务执行

lucataco · Qwen2 72B

A
厂商:Qwen / Alibaba 模型:Qwen2 72B 来源平台:github 最后复核:2026-06-27T08:07:03Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

lucataco 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:07:03Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The lucataco GitHub project implements Qwen/Qwen2-72B-Instruct-GPTQ-Int4 as a Cog model, enabling local Cog prediction and publishing to Replicate-style model serving workflows.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public README documents a working Cog wrapper and a sample prediction command using Qwen2 72B Instruct GPTQ-Int4 for short text generation.

模型作用:Qwen2 72B 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2 72B Instruct GPTQ-Int4 is the generation engine wrapped by the Cog project; the repository supplies deployment glue around the model.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:GitHub HTML returned a transient 502 from this environment, so the reachable raw README is used as both evidence and artifact; the README links to the exact Qwen/Qwen2-72B-Instruct-GPTQ-Int4 model.

原始记录:lucataco packaged Qwen2 72B Instruct GPTQ-Int4 as a Cog model

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

lalopenguin 使用 Qwen3 235B A22B 2507 处理真实任务执行

lalopenguin · Qwen3 235B A22B 2507

A
厂商:Qwen / Alibaba 模型:Qwen3 235B A22B 2507 来源平台:huggingface_spaces 最后复核:2026-06-27T08:01:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

lalopenguin 公开的真实任务执行案例,来源为 huggingface_spaces,复核于 2026-06-27T08:01:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上发布一个需要 Hugging Face 登录的 Gradio Blocks 应用,展示由 fireworks-ai API 服务的 Qwen/Qwen3-235B-A22B-Instruct-2507 模型。

公开产物:公开 app.py 明确写有 This Space showcases the Qwen/Qwen3-235B-A22B-Instruct-2507 model,并通过 gr.load("models/Qwen/Qwen3-235B-A22B-Instruct-2507", provider="fireworks-ai") 加载模型产物。

模型作用:Qwen3-235B-A22B-Instruct-2507 作为 Gradio demo 的被调用模型,为用户交互提供文本生成/聊天推理能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:社区 Hugging Face demo,需登录使用 Inference API;证据为公开 Space 源码和可访问产物页。

原始记录:lalopenguin 在 Hugging Face Space 发布 Qwen3-235B-A22B-Instruct-2507 推理 Demo

已有真实案例 真实任务执行huggingface_spacesA 类可核验real_case auto_approved 进入模型卡精选

Neural Magic / Red Hat AI 使用 Qwen2 72B 处理智能体流程编排

Neural Magic / Red Hat AI · Qwen2 72B

A
厂商:Qwen / Alibaba 模型:Qwen2 72B 来源平台:huggingface 最后复核:2026-06-27T08:07:03Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Neural Magic / Red Hat AI 公开的智能体工作流案例,来源为 huggingface,复核于 2026-06-27T08:07:03Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Neural Magic quantized Qwen2 72B Instruct weights and activations to FP8 and published a Red Hat AI Hugging Face artifact intended for assistant-like chat deployment with vLLM.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public model card states the optimized artifact is ready for inference with vLLM >= 0.5.0 and reduces disk size and GPU memory requirements by approximately 50% compared with the original precision.

模型作用:Qwen2 72B 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Qwen2 72B Instruct provides the base instruction-following model; the case demonstrates its adaptation into a lower-memory FP8 deployment artifact.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a concrete public derivative/deployment artifact, not an end-customer story; evidence explicitly says it is a quantized version of Qwen2-72B-Instruct.

原始记录:Neural Magic / Red Hat AI created an FP8 vLLM-ready Qwen2 72B Instruct artifact

已有真实案例 智能体工作流huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

UpThink team (Upstage AI Ambassador… 使用 Solar Pro 2 处理知识检索和问答

UpThink team (Upstage AI Ambassador project) · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:github 最后复核:2026-06-27T08:06:43Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

UpThink team (Upstage AI Ambassador project) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T08:06:43Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The UpThink project processes images embedded in Obsidian markdown notes: it extracts text and document structure with Upstage Document Parse, then sends the extracted text to Solar Pro 2 to create about 50 words of Kor…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Generated alt text is automatically inserted below image links in the markdown note, reducing manual conversion of visual information into accessible text for personal knowledge management.

模型作用:Solar Pro 2 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 2 performs the language-generation step that turns OCR/document-parse output into concise Korean image descriptions; the public code defaults the generator model to solar-pro2 and calls the Upstage-compatible …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public GitHub project evidence; use case is an Upstage AI Ambassador/demo project rather than a named enterprise deployment, but it includes reachable code and README documentation binding the feature to Solar Pro 2.

原始记录:UpThink uses Solar Pro 2 to generate Korean alt text for Obsidian note images

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

UpThink team (Upstage AI Ambassador… 使用 Solar Pro 2 处理知识检索和问答

UpThink team (Upstage AI Ambassador project) · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:github 最后复核:2026-06-27T08:06:43Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

UpThink team (Upstage AI Ambassador project) 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T08:06:43Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:UpThink analyzes markdown note content and a user-defined tagging guideline, then asks Solar Pro 2 to generate new tags that can be compared against existing Obsidian vault tags and inserted into YAML frontmatter.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The system produces a parsed list of tags, aligns them with existing tag conventions through embedding-based comparison, and writes the final tags back into the note frontmatter.

模型作用:Solar Pro 2 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 2 is the LLM used to infer relevant tags from the note body and guideline prompt; the public TagGenerator class defaults to model solar-pro2 and sends system/user messages to Upstage's API.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public GitHub project evidence; same broader UpThink artifact as the alt-text case, but this is a separate documented feature with a distinct implementation file and task output.

原始记录:UpThink uses Solar Pro 2 to recommend consistent tags for Obsidian markdown notes

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ShenSeanChen 使用 Kimi K2 0905 处理研究分析和报告生成

ShenSeanChen · Kimi K2 0905

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 0905 来源平台:github 最后复核:2026-06-27T08:00:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ShenSeanChen 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:00:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建 FastAPI + LangGraph 的 Deep Research Agent 后端,让用户可在 OpenAI、Anthropic 和 Kimi K2 间选择模型并通过流式 API 执行研究任务。

公开产物:公开仓库提供可运行后端、Docker/GCP Cloud Run 部署说明、模型对比接口和 Kimi K2 集成指南;KIMI_K2_INTEGRATION.md 明确记录 Kimi K2 0905 preview 的 128k token 配置与 Moonshot API 路由方式。

模型作用:Kimi K2 0905 被作为 deep research 工作流的可选推理模型,通过 Moonshot API/Anthropic 兼容格式参与问题澄清、研究执行、报告生成和多模型比较。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 仓库与集成文档;仓库也含教学视频链接,但代码产物和模型绑定可直接核验。

原始记录:ShenSeanChen 用 Kimi K2 0905 集成 Deep Research Agent 后端

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Team5_NLP_Upstage project team 使用 Solar Pro 2 处理代码审查和测试生成

Team5_NLP_Upstage project team · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:github 最后复核:2026-06-27T08:06:43Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Team5_NLP_Upstage project team 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:06:43Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Team5_NLP_Upstage term project uses Solar Pro 2 with RAG, domain routing, prompt engineering, table parsing, and majority-vote ensembling to answer EWHA and MMLU-style questions from supplied datasets.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository reports 88% total performance on 5_final.csv, including 100% on EWHA questions and 76% on MMLU questions, and provides runnable code that writes answers to 5_final.csv.

模型作用:Solar Pro 2 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 2 is instantiated through LangChain ChatUpstage for both EWHA voting and MMLU domain-specific answering, while Upstage embeddings retrieve context; the model generates the final answer choices after routing an…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Academic project with evaluation-style output, not a commercial deployment; accepted because the repo contains a concrete runnable QA application with public code, task, output, and explicit solar-pro2 model usage rathe…

原始记录:Team5_NLP_Upstage built a RAG and prompt-routing QA pipeline with Solar Pro 2

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Coobiw 使用 Kimi K2 0905 处理软件工程任务执行

Coobiw · Kimi K2 0905

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 0905 来源平台:github 最后复核:2026-06-27T08:00:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Coobiw 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:00:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发一个抓取小红书每日热搜、保存 JSON/CSV 数据并生成响应式 HTML 深度分析报告的 Python 项目。

公开产物:公开仓库包含爬虫脚本、配置文件、分析与报告生成模块、任务记录以及示例输出目录;README 写明项目由 Claude Code + Kimi-K2-0905 在 10-20 分钟内完成。

模型作用:Kimi K2 0905 作为编码/开发辅助模型参与项目实现,贡献了爬虫、数据分析、报告生成和项目结构搭建相关代码产物。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该案例是模型辅助开发公开项目,而非线上 SaaS 生产部署;仓库 README 明确绑定 Kimi-K2-0905 和具体产物。

原始记录:Coobiw 用 Kimi K2 0905 辅助完成小红书热搜爬虫与分析报告项目

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

VaishnaviAnalyzed 使用 Kimi K2 0905 处理真实任务执行

VaishnaviAnalyzed · Kimi K2 0905

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 0905 来源平台:github 最后复核:2026-06-27T08:00:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

VaishnaviAnalyzed 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:00:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发一个 Streamlit 聊天机器人界面,支持会话历史、实时 token streaming,并通过 Groq API 调用 Kimi K2 Instruct 0905。

公开产物:公开仓库包含 README、kimiapp.py 和 requirements.txt;kimiapp.py 中明确使用 model="moonshotai/kimi-k2-instruct-0905" 并实现 st.chat_input、st.session_state 和流式输出。

模型作用:Kimi K2 Instruct 0905 是聊天机器人的核心回复生成模型,负责根据用户输入和会话上下文生成实时流式助手响应。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自代码文件和 README;这是开源演示应用,未声明大规模生产使用。

原始记录:VaishnaviAnalyzed 用 Kimi K2 Instruct 0905 构建 Streamlit 聊天机器人

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Arfa-Ahsan 使用 Kimi K2 0905 处理智能体流程编排

Arfa-Ahsan · Kimi K2 0905

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 0905 来源平台:github 最后复核:2026-06-27T08:00:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Arfa-Ahsan 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:00:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建一个 Streamlit 邮件内容优化工具,基于 LangGraph/Reflect Evaluator Pattern 生成、评估并迭代改进专业邮件。

公开产物:公开仓库包含 Streamlit 前端、email_optimizer.py 工作流和 README;代码中 ChatGroq 明确配置 model="moonshotai/kimi-k2-instruct-0905",README 描述可生成初稿、给出反馈并输出优化邮件。

模型作用:Kimi K2 Instruct 0905 同时承担邮件生成与结构化反馈评估,驱动反思式迭代流程优化语气、清晰度、专业性、语法和行动性。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开代码与 README;属于可运行开源应用案例,未提供外部部署地址。

原始记录:Arfa-Ahsan 用 Kimi K2 Instruct 0905 构建邮件内容优化 Agent

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

yousefhafed2222-gif 使用 Kimi K2 0905 处理智能体流程编排

yousefhafed2222-gif · Kimi K2 0905

A
厂商:Kimi / Moonshot AI 模型:Kimi K2 0905 来源平台:github 最后复核:2026-06-27T08:00:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

yousefhafed2222-gif 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:00:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:创建一个 Node.js/Express 的 OpenClaw-style AI Agent 服务,提供 /chat 接口并通过 Groq OpenAI-compatible chat completions 调用 Kimi K2 0905。

公开产物:公开仓库包含 index.js、package.json 和 README;index.js 明确在 /chat 路由中调用 model="moonshotai/kimi-k2-instruct-0905",根路径返回 OpenClaw Agent Running。

模型作用:Kimi K2 Instruct 0905 是该 Agent 服务的对话推理后端,接收用户 message 并生成 API 响应。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开代码;README 较简短,案例规模较小,但模型、使用者、任务和可访问代码产物均可核验。

原始记录:yousefhafed2222-gif 用 Kimi K2 0905 接入 OpenClaw 风格聊天 Agent

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Qwen / Alibaba Cloud 使用 Qwen3 Max Thinking (Preview) 处理知识检索和问答

Qwen / Alibaba Cloud · Qwen3 Max Thinking (Preview)

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking (Preview) 来源平台:official_web 最后复核:2026-06-27T08:00:32Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Qwen / Alibaba Cloud 公开的知识库与检索问答案例,来源为 官方页面,复核于 2026-06-27T08:00:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Alibaba Cloud states that Qwen3-Max-Thinking is available in Qwen Chat, where users can interact with the model and its adaptive tool-use capabilities including Search, Memory and Code Interpreter.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Qwen Chat entry point is available for interactive conversations with Qwen3-Max-Thinking, and the official post documents the model's tool-use and reasoning behavior for chat users.

模型作用:Qwen3 Max Thinking (Preview) 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-Max-Thinking provides the reasoning layer for Qwen Chat, autonomously selecting retrieval, memory and code-interpreter tools during conversations to improve factuality, personalization and computational reasoning.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Official vendor-operated product case; evidence is stronger for public availability and tool-use positioning than for a named third-party customer workflow.

原始记录:Qwen Chat exposes Qwen3-Max-Thinking with adaptive Search, Memory and Code Interpreter tools

已有真实案例 知识库与检索问答官方页面A 类可核验real_case auto_approved 进入模型卡精选

OpenRouter 使用 Qwen3 Max Thinking (Preview) 处理智能体流程编排

OpenRouter · Qwen3 Max Thinking (Preview)

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking (Preview) 来源平台:official_web 最后复核:2026-06-27T08:00:32Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

OpenRouter 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T08:00:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenRouter publishes Qwen: Qwen3 Max Thinking under model slug qwen/qwen3-max-thinking with pricing, context length, provider routing and quick-start/playground links for developers.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Developers can access Qwen3-Max-Thinking through OpenRouter's model page and API; the page shows a 262K context window, model pricing, released date and Alibaba Cloud International as the hosting provider.

模型作用:Qwen3 Max Thinking (Preview) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-Max-Thinking supplies the high-stakes reasoning model behind OpenRouter's API listing, enabling developers to route reasoning-heavy chat and agent requests through a standard OpenRouter interface.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Platform/API distribution case rather than a downstream application case; page is public and directly names the model, provider, pricing and access artifact.

原始记录:OpenRouter lists Qwen3-Max-Thinking as a public API model routed to Alibaba Cloud International

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

DeepInfra 使用 Qwen3 Max Thinking (Preview) 处理智能体流程编排

DeepInfra · Qwen3 Max Thinking (Preview)

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking (Preview) 来源平台:official_web 最后复核:2026-06-27T08:00:32Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

DeepInfra 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T08:00:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DeepInfra hosts a public demo/product page for Qwen/Qwen3-Max-Thinking, presenting it as Qwen's flagship reasoning model and exposing the model as an interactive hosted artifact.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public DeepInfra page is reachable and displays Qwen/Qwen3-Max-Thinking as a demo with model description, reasoning and tool-use capabilities, and benchmark/context information for prospective users.

模型作用:Qwen3 Max Thinking (Preview) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-Max-Thinking is the model being served by DeepInfra for public experimentation, contributing long-context reasoning, instruction following, agentic behavior and adaptive tool-use capabilities to the hosted demo/AP…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a provider-hosted demo/product page, not a named customer story; still meets public artifact and explicit model binding requirements.

原始记录:DeepInfra provides a public Qwen/Qwen3-Max-Thinking demo page

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

haimaker.ai 使用 Qwen3 Max Thinking (Preview) 处理文档理解和结构化处理

haimaker.ai · Qwen3 Max Thinking (Preview)

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking (Preview) 来源平台:official_web 最后复核:2026-06-27T08:00:32Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

haimaker.ai 公开的文档理解与处理案例,来源为 官方页面,复核于 2026-06-27T08:00:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:haimaker.ai publishes a model page for qwen/qwen3-max-thinking and documents its OpenAI-compatible chat completions endpoint for API users.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public haimaker.ai artifact lists Qwen3 Max Thinking with 262,144-token context, 32,768 max output tokens, pricing, function calling and reasoning support, plus API usage instructions for https://api.haimaker.ai/v1/…

模型作用:Qwen3 Max Thinking (Preview) 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-Max-Thinking provides haimaker.ai users with a reasoning-capable chat model that supports function calling and long-context completions through a standard OpenAI-compatible API.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Provider distribution/use case; evidence is strong for public model serving and API artifact, but not a downstream customer deployment.

原始记录:haimaker.ai offers qwen/qwen3-max-thinking through an OpenAI-compatible API

已有真实案例 文档理解与处理官方页面A 类可核验real_case auto_approved 进入模型卡精选

Gizem Ergün / Hochschule Hannover 使用 Grok 3 mini Reasoning (high) 处理代码审查和测试生成

Gizem Ergün / Hochschule Hannover · Grok 3 mini Reasoning (high)

A
厂商:xAI / Grok 模型:Grok 3 mini Reasoning (high) 来源平台:github 最后复核:2026-06-27T08:13:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Gizem Ergün / Hochschule Hannover 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:13:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Run aspect-based sentiment analysis on German Google reviews of stationary nursing homes, comparing prompting methods including Few-Shot, QAIE, and Syn-Chain prompting.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository reports Grok 3 Mini results for the task, including 81.22% accuracy / 0.7842 Macro-F1 with Few-Shot prompting and 82.45% accuracy / 0.7953 Macro-F1 with Syn-Chain prompting.

模型作用:Grok 3 mini Reasoning (high) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Grok 3 Mini via the paid xAI API supplied the text classification and multi-step Syn-Chain reasoning used to extract aspects/opinions and assign sentiment labels from the review text.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence names Grok 3 Mini rather than an API parameter spelling of reasoning_effort=high; mapped to the Model Atlas Grok 3 Mini reasoning-high variant because the task uses the same xAI Grok 3 Mini reasoning-capable mo…

原始记录:Hochschule Hannover bachelor project used Grok 3 Mini for aspect-based sentiment analysis of German nursing-home reviews

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Suiundukov Mederbek / Ala-Too Inter… 使用 Grok 3 mini Reasoning (high) 处理软件工程任务执行

Suiundukov Mederbek / Ala-Too International University · Grok 3 mini Reasoning (high)

A
厂商:xAI / Grok 模型:Grok 3 mini Reasoning (high) 来源平台:github 最后复核:2026-06-27T08:13:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Suiundukov Mederbek / Ala-Too International University 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:13:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Provide a Streamlit chat interface where non-technical users ask natural-language questions over a normalized TiDB Serverless database containing the Olist Brazilian e-commerce dataset.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The technical report describes a working end-to-end app that inspects schema, drafts SQL, checks and executes queries, synthesizes answers, and renders automatic charts for business questions such as category revenue, o…

模型作用:Grok 3 mini Reasoning (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok-3-mini, accessed through the xAI API with LangChain ChatOpenAI, powers the ReAct SQL agent loop: schema inspection, SQL generation, query checking, retry/refinement, and natural-language answer synthesis.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The technical report explicitly says ChatOpenAI uses Grok-3-mini via https://api.x.ai/v1; README/app comments contain an inconsistent Groq/Llama mention, so the report is used as the binding evidence. The reasoning-high…

原始记录:Ala-Too internship project built a Grok-powered Text-to-SQL agent for Olist e-commerce analytics

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Ayush / 4yu5h-crtl 使用 Grok 3 mini Reasoning (high) 处理真实任务执行

Ayush / 4yu5h-crtl · Grok 3 mini Reasoning (high)

A
厂商:xAI / Grok 模型:Grok 3 mini Reasoning (high) 来源平台:github 最后复核:2026-06-27T08:13:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Ayush / 4yu5h-crtl 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:13:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Convert a user prompt into a complete, ready-to-use HTML website inside a Streamlit web editor, with live editing and deployment options for GitHub Pages, Netlify, Vercel, or local export.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The application calls OpenRouter with model x-ai/grok-3-mini-beta, extracts the model response as HTML, inserts it into the editor state, and lets the user save, remix, or deploy the generated site.

模型作用:Grok 3 mini Reasoning (high) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 3 Mini Beta generates the full HTML/CSS/JavaScript page content and can stream incremental HTML chunks back to the editor for AI-assisted frontend development.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence uses OpenRouter model id x-ai/grok-3-mini-beta and README wording 'Grok 3 Mini Beta'; mapped to Grok 3 Mini reasoning-high family because the public artifact does not expose the private reasoning_effort setting.

原始记录:Frontend_Builder uses Grok 3 Mini Beta to generate deployable HTML websites from prompts

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mvanhorn / OpenClaw skill ecosystem 使用 Grok 3 mini Reasoning (high) 处理软件工程任务执行

mvanhorn / OpenClaw skill ecosystem · Grok 3 mini Reasoning (high)

A
厂商:xAI / Grok 模型:Grok 3 mini Reasoning (high) 来源平台:github 最后复核:2026-06-27T08:13:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mvanhorn / OpenClaw skill ecosystem 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:13:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Expose xAI Grok capabilities as an installable OpenClaw skill for chat, vision analysis, web/X search, responses API tools, model comparison, code generation, batch processing, and step-by-step reasoning prompts.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a working Node.js skill and CLI scripts such as chat.js, models.js, and batch.js; the README lists Grok-3-mini as a supported model with 'reasoning effort control' and includes example commands f…

模型作用:Grok 3 mini Reasoning (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 3 Mini supplies fast text chat and reasoning-capable responses inside the OpenClaw skill, enabling users to ask Grok questions, run step-by-step problem solving, compare models, and integrate xAI calls into agent w…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an integration/tooling artifact rather than an external customer story. It explicitly identifies Grok-3-mini and reasoning effort control, but the public README does not show a captured reasoning_effort=high API…

原始记录:OpenClaw xAI skill added Grok 3 Mini chat and reasoning workflows for agent users

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Simon Pierre Boucher 使用 Grok 3 mini Reasoning (high) 处理软件工程任务执行

Simon Pierre Boucher · Grok 3 mini Reasoning (high)

A
厂商:xAI / Grok 模型:Grok 3 mini Reasoning (high) 来源平台:github 最后复核:2026-06-27T08:13:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Simon Pierre Boucher 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:13:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Provide a Python command-line chat agent for Grok models with persistent conversations, per-agent YAML configuration, file injection, streaming/non-streaming responses, search over conversation history, and JSON/TXT/Mar…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository defines Grok 3 Mini support through the grok3mini alias mapping to grok-3-mini-latest, documents a multi-model CLI workflow, and ships code for persisted sessions, exports, logging, and API calls to https…

模型作用:Grok 3 mini Reasoning (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 3 Mini is one of the selectable xAI inference backends used to generate streaming chat responses for fast, cost-efficient CLI assistant sessions with local history and exportable transcripts.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence binds the artifact to grok-3-mini-latest rather than a visible high-effort setting; included as a public, working Grok 3 Mini use case mapped to the Model Atlas reasoning-high entry by model family.

原始记录:Grok CLI Agent uses Grok 3 Mini for persistent multi-session command-line assistants

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

HPC-AI Tech / ColossalAI 使用 Grok-1 处理真实任务执行

HPC-AI Tech / ColossalAI · Grok-1

A
厂商:xAI / Grok 模型:Grok-1 来源平台:official_web 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

HPC-AI Tech / ColossalAI 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:HPC-AI Tech 针对 xAI Grok-1 314B 开源权重,构建 PyTorch + Hugging Face 版本的推理示例,并用 ColossalAI 张量并行在多 GPU 环境中运行生成推理。

公开产物:公开了 ColossalAI 的 Grok-1 推理代码、运行脚本和 hpcai-tech/grok-1 Hugging Face 权重;项目 README 明确称 Grok-1 推理加速 3.8x,并给出 8x A100 80GB 运行方式。

模型作用:Grok-1 是被转换和运行的目标大模型;案例围绕其 314B MoE 权重展开,ColossalAI 的贡献是让该模型可通过 PyTorch/Hugging Face 接口加载并进行加速推理。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是基础设施/模型适配案例,不是终端业务部署;需要多 GPU 硬件,公开证据主要来自 HPC-AI 官方博客、GitHub 示例和 Hugging Face 模型仓库。

原始记录:HPC-AI Tech 用 ColossalAI 将 Grok-1 转为 PyTorch/Hugging Face 推理并加速 3.8x

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

ggml-org / llama.cpp contributors 使用 Grok-1 处理真实任务执行

ggml-org / llama.cpp contributors · Grok-1

A
厂商:xAI / Grok 模型:Grok-1 来源平台:github 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ggml-org / llama.cpp contributors 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:llama.cpp 贡献者为 Grok-1 增加模型架构、GGUF 常量、tensor mapping 和转换脚本支持,使 Grok-1 能进入 llama.cpp/ggml 的本地推理生态。

公开产物:公开 PR/patch 标题为“Add support for Grok model architecture”,变更覆盖 convert-hf-to-gguf.py、gguf constants、tensor_mapping 和 llama.cpp;后续 Grok-1 GGUF 仓库明确依赖该 PR 的兼容性。

模型作用:Grok-1 是该工程改造的明确目标模型;它的架构和权重格式需求驱动了 llama.cpp 对 Grok 模型类型、张量映射和 GGUF 转换路径的新增支持。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是开源运行时支持案例而非业务产品;PR 页面可公开访问,patch 可直接核验新增 Grok 架构支持。

原始记录:ggml-org / llama.cpp 为 Grok-1 增加本地 GGUF 推理架构支持

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Arki05 使用 Grok-1 处理真实任务执行

Arki05 · Grok-1

A
厂商:xAI / Grok 模型:Grok-1 来源平台:huggingface 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Arki05 公开的真实任务执行案例,来源为 huggingface,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Arki05 将 Grok-1 开源权重制作成多个 GGUF 量化版本,并提供 llama.cpp server 的直接加载命令,降低本地部署 Grok-1 的存储和运行门槛。

公开产物:公开 Hugging Face 仓库列出 Q2_K、IQ3_XS 等 Grok-1 分片 GGUF 文件和运行命令;README 明确说明兼容 llama.cpp 的 Grok-1 support PR。

模型作用:Grok-1 是量化与分片的源模型;案例产物是围绕 Grok-1 权重生成的 GGUF 文件,使社区可在 llama.cpp 路径上尝试加载该模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是社区模型打包/部署案例;不是官方 xAI 发布物,量化质量和运行效果需由使用者自行验证。

原始记录:Arki05 发布 Grok-1 GGUF 量化包用于 llama.cpp 本地加载

已有真实案例 真实任务执行huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

keyfan 使用 Grok-1 处理软件工程任务执行

keyfan · Grok-1

A
厂商:xAI / Grok 模型:Grok-1 来源平台:huggingface 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

keyfan 公开的代码代理与软件工程案例,来源为 huggingface,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:keyfan 基于 xAI Grok-1 原始权重和 grok-1 仓库脚本,制作非官方 dequantized Hugging Face Transformers 格式权重,便于用 Transformers 工具链加载。

公开产物:公开了 keyfan/grok-1-hf 模型仓库;README 明确写明这是 grok-1 的非官方 dequantized HF Transformers 格式权重,并链接转换脚本和原始 grok-1 repo。

模型作用:Grok-1 是被转换的模型本体;该案例的价值来自将 Grok-1 权重映射到 Transformers 生态,使后续推理、评测或二次工程更容易接入。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是社区转换案例,README 包含基准结果但候选只采用其权重转换和公开仓库产物作为证据;非官方维护。

原始记录:keyfan 将 Grok-1 开源权重转换为 Hugging Face Transformers 格式

已有真实案例 代码代理与软件工程huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

Omnibus 使用 Grok-1 处理真实任务执行

Omnibus · Grok-1

A
厂商:xAI / Grok 模型:Grok-1 来源平台:huggingface 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Omnibus 公开的真实任务执行案例,来源为 huggingface,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Omnibus 创建 Gradio Space,默认模型字段为 xai-org/grok-1,通过 Hugging Face Hub API 搜索模型并尝试 gr.load 加载模型,输出 Gradio 加载信息和 Hub 模型详情。

公开产物:公开 Space“Grok 1 Test”和 app.py;代码中 models 列表固定包含 xai-org/grok-1,并在页面加载或点击按钮时返回模型加载错误/详情 JSON,形成可访问的 Grok-1 检查 Demo。

模型作用:Grok-1 是 Demo 默认检查和加载的目标模型;应用围绕 xai-org/grok-1 的 Hub 可发现性和 Gradio 加载行为构建。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是轻量公开 Demo,功能是模型加载/元数据检查,不代表完整生产部署;证据来自公开 Space 源码且精确绑定 xai-org/grok-1。

原始记录:Omnibus 在 Hugging Face Space 构建 Grok-1 模型加载与 Hub 信息检查 Demo

已有真实案例 真实任务执行huggingfaceA 类可核验real_case auto_approved 进入模型卡精选

MemeCalculate / 魔因漫创 Moyin Creator 使用 Seedance 2.0 处理智能体流程编排

MemeCalculate / 魔因漫创 Moyin Creator · Seedance 2.0

A
厂商:ByteDance Seed 模型:Seedance 2.0 来源平台:github 最后复核:2026-06-27T08:22:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MemeCalculate / 魔因漫创 Moyin Creator 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:22:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开源 AI 影视生产工具把剧本、角色、场景、导演分镜串成批量生产链路,并在 S 级板块调用 Seedance 2.0 做多模态视频创作。

公开产物:仓库公开了可运行桌面/源码项目、工作流教程和多张界面截图;README 明确列出 Seedance 2.0 多镜头合并叙事视频生成、首帧图拼接、批量生视频等功能。

模型作用:Seedance 2.0 负责把分镜、角色/场景参考图、视频/音频引用和镜头提示词转成连贯叙事视频,是该工具从分镜到成片环节的核心视频生成模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README 自述和公开仓库,未独立验证生成视频质量;仓库与说明页面可访问。

原始记录:魔因漫创用 Seedance 2.0 批量化生成短剧/动漫视频

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

serithemage 使用 Solar Mini 处理软件工程任务执行

serithemage · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:GitHub 最后复核:2026-06-27T08:25:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

serithemage 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:25:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an MCP server that lets Claude Code or Claude Desktop send chat-completion requests to Upstage Solar models, including selecting solar-mini for fast and efficient responses.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A published open-source Node.js MCP server with a solar_chat tool, Claude Code/Desktop setup instructions, usage accounting resources, and a model list that explicitly includes solar-mini.

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini is one of the callable chat models behind the MCP tool; it generates the assistant response for user-supplied messages when the model parameter is set to solar-mini.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository README and source; default model is solar-pro2, but solar-mini is explicitly supported and selectable.

原始记录:solar-mcp exposes Solar Mini as an MCP chat tool for Claude clients

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Hyeongseob91 使用 Solar Pro 3 处理软件工程任务执行

Hyeongseob91 · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:GitHub 最后复核:2026-06-27T08:18:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Hyeongseob91 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:18:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Streamlit dashboard that crawls Moltbook posts and uses Upstage Solar Pro 3 to classify topics, writing patterns, agent personas, sentiment, trends and meme signals in AI-agent community discourse.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a runnable dashboard with real-time crawling, background analysis, SQLite caching, four analysis tabs, export features and an analysis schema for Solar Pro 3 outputs.

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 3 is the LLM analysis engine for comprehensive post analysis, producing structured labels for topic, persona, sentiment and trend fields that drive the dashboard views.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-declared in a public GitHub README; it is an open-source project rather than independently verified production deployment.

原始记录:Agent Genome Watcher analyzes Moltbook AI-agent community discourse with Solar Pro 3

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zzdx713 使用 Seedance 2.0 处理多模态内容处理

zzdx713 · Seedance 2.0

A
厂商:ByteDance Seed 模型:Seedance 2.0 来源平台:github 最后复核:2026-06-27T08:22:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zzdx713 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:22:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建面向内容创作者、设计师和营销人员的 Seedance 2.0 Web 应用,用户上传 1-5 张参考图片并输入自然语言描述后生成 AI 视频。

公开产物:公开仓库包含 React/Express 应用、Docker 部署方式、即梦 SessionID 配置说明、生成流程截图,以及 Seedance 2.0/Seedance 2.0 Fast 双模型选择、实时进度、视频预览和下载功能说明。

模型作用:Seedance 2.0 通过即梦接口执行核心图/文到视频生成,支持多图参考、首帧/尾帧/全能参考、4-15 秒时长与多种画幅输出。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 中列出的体验站 seedance2.duckcloud.fun 在本次 DNS 检查中不可解析,因此 artifact_url 使用可访问的 GitHub 仓库;证据为项目自述。

原始记录:zzdx713 开源 Seedance 2.0 参考图视频生成 Web 应用

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Hongduk 使用 Solar Pro 3 处理知识检索和问答

Hongduk · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:GitHub 最后复核:2026-06-27T08:18:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Hongduk 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T08:18:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a personal Korean exam-study tutor for a middle-school student by parsing scanned textbook PDFs with Upstage Document Parse and loading each unit directly into Solar Pro 3's long context for five study modes.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The app offers unit overview, key concept review, step-by-step interactive tutoring, quiz generation with grading/explanations and mock exams through a Streamlit interface.

模型作用:Solar Pro 3 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 3 is the default tutoring LLM; its 128K context window is used instead of RAG to keep an entire unit in context and generate Korean study explanations, questions, grading and feedback.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-declared in a public GitHub README; textbook files are excluded for copyright, but the repository describes the workflow and runnable app structure.

原始记录:Korean middle-school personal tutor app uses Solar Pro 3 long context over textbook units

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

UpstageAI 使用 Solar Mini 处理文档理解和结构化处理

UpstageAI · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:GitHub 最后复核:2026-06-27T08:25:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

UpstageAI 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T08:25:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Provide n8n community nodes so workflow builders can call Upstage Solar chat completions inside visual automations and set the chat model to solar-mini.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An installable n8n community-node package with a Solar Chat Model node, API-key credential flow, UI installation instructions, and supported-model documentation listing solar-mini, solar-pro, and solar-pro2.

模型作用:Solar Mini 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini supplies fast chat-completion output for the n8n Solar Chat Model node when users choose the solar-mini option in workflow configuration.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public repo evidence binds the model explicitly; this is a developer integration artifact rather than a named end-customer deployment.

原始记录:UpstageAI n8n-nodes-solar adds Solar Mini chat completions to n8n workflows

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chongdashu 使用 Seedance 2.0 处理可玩交互原型构建

chongdashu · Seedance 2.0

A
厂商:ByteDance Seed 模型:Seedance 2.0 来源平台:github 最后复核:2026-06-27T08:22:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chongdashu 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T08:22:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:为游戏开发生成一致角色的 idle、walk、attack 动画和运行时 spritesheet,使用图像生成锚点后通过 Seedance 2.0 的 image-to-video 步骤生成运动循环。

公开产物:仓库公开了提示词模板、参考图、walk/idle/attack GIF 预览、接触表和最终 runtime spritesheet,并在 README 中声明这些动画由该管线生成。

模型作用:Seedance 2.0 负责将角色锚点图转成走路/攻击/待机动画片段,为后续背景移除、帧归一化和游戏运行时精灵表制作提供运动素材。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开仓库、预览素材和作者 README;模型调用细节为作者自述,仓库产物可访问。

原始记录:chongdashu 用 Seedance 2.0 管线生成游戏角色动画精灵表

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

gem-squared 使用 Solar Pro 3 处理软件工程任务执行

gem-squared · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:GitHub 最后复核:2026-06-27T08:18:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

gem-squared 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:18:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Implement a runtime judgment verification layer for Korean health-insurance claim documents, combining Upstage Document AI extraction with Solar Pro 3 pre-check, judgment and evidence-audit stages.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public demo and repository include three preloaded scenarios: a clean claim, a trap claim where verification catches an unsupported rider certificate claim, and a malicious-injection claim blocked before LLM process…

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 3 performs L1 precondition validation, the core Korean-language adjudication over extracted facts, and L2 epistemic auditing that traces cited evidence references back to source anchors.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-declared by the public repository and demo; it is a demo governance layer, not a named insurer production case.

原始记录:Solar Trust Gate uses Solar Pro 3 to adjudicate and audit Korean insurance claims

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

UpstageAI 使用 Solar Mini 处理软件工程任务执行

UpstageAI · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:GitHub 最后复核:2026-06-27T08:25:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

UpstageAI 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:25:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Package Upstage Solar chat models for n8n workflows, including AI Agent-compatible language-model nodes, chat-completion nodes, embeddings, and document-processing operations.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public n8n community-node repository that documents Solar Chat Models with solar-mini support and implements Upstage Solar Chat / language-model nodes for use in n8n workflow automation.

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini is an explicitly supported model option and default chat model in the node implementation, producing conversational completions used by n8n workflows and AI-agent chains.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public UpstageAI integration artifact; overlaps with n8n-nodes-solar but is a separate repository/package with broader Upstage node coverage.

原始记录:UpstageAI n8n-nodes-upstage uses Solar Mini for n8n AI-agent and automation nodes

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SamurAIGPT 使用 Seedance 2.0 处理多模态内容处理

SamurAIGPT · Seedance 2.0

A
厂商:ByteDance Seed 模型:Seedance 2.0 来源平台:github 最后复核:2026-06-27T08:22:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SamurAIGPT 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:22:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建一个 Next.js AI 视频生成、编辑和管理工作台,提供登录、计费/积分、图片持久化和异步视频生成,用于浏览器内生成 Seedance 2.0 视频。

公开产物:README 公开了 Seedance 2 Generator 项目、Seedance 2.0 与 Seedance 2 Mini 模型入口、Muapi playground 链接和线上 Vercel 应用;本次核验线上应用与 GitHub 仓库均返回 200。

模型作用:Seedance 2.0 是该工作台的核心视频生成引擎之一,负责 text-to-video 和 image-to-video 任务,应用层围绕其生成任务做认证、额度、持久化和异步管理。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目 README 同时提到 Seedance 2 Mini 和未来版本;本条仅采纳 README 明确写到的 Seedance 2.0 使用,未登录线上应用验证实际生成。

原始记录:SamurAIGPT 开源 Seedance 2 Generator 视频生成 SaaS 工作台

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Mixup-Agent / solsol team 使用 Solar Pro 3 处理软件工程任务执行

Mixup-Agent / solsol team · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:GitHub 最后复核:2026-06-27T08:18:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Mixup-Agent / solsol team 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:18:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a web-based voice mock-interview service that analyzes resumes, cover letters, portfolios, company names, job postings and company news, then runs multiple interview agents tailored to a target employer's intervie…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository describes a LangGraph multi-agent service with company-style inference, resume/trend/stress interviewers, answer-quality gating, STT/TTS, source-tracked trend questions and final feedback reports.

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro3 is used for the meta router and judge at high reasoning effort, and for resume, trend and stress interview agents at low reasoning effort; prompt caching and Solar Mini are used for cost/speed tradeoffs.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-declared in a hackathon/open-source GitHub README; it demonstrates a concrete app but not independently verified commercial production usage.

原始记录:solsol builds a voice-based Korean mock interview service with Solar Pro3 multi-agent routing

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

thisiskorea 使用 Solar Mini 处理软件工程任务执行

thisiskorea · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:GitHub 最后复核:2026-06-27T08:25:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

thisiskorea 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:25:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a Korean-law question-answering notebook that retrieves relevant law PDFs and asks Upstage chat to produce legal answers grounded in the retrieved legal material.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public repository with Korean law PDF artifacts and a notebook that extracts PDF text, indexes law documents, retrieves relevant files, and returns legal-answer sections covering facts, applicable law, application, an…

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The repository title and README identify the project as an Upstage Solar Mini law application; the notebook uses langchain_upstage ChatUpstage to generate legal answers from retrieved law text.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The notebook uses ChatUpstage without an explicit model parameter in the visible code, so the binding relies on the repository name/README rather than an in-code solar-mini literal; repository also contains an exposed A…

原始记录:upstage_solar_mini_law answers Korean legal questions over law PDFs

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Sioonn 使用 Solar Pro 3 处理软件工程任务执行

Sioonn · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:GitHub 最后复核:2026-06-27T08:18:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Sioonn 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:18:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Improve Korean natural-language fashion image search by generating user-style Korean search queries from product captions and using them to align a SigLIP 2 text tower with fashion product images.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project reports a lightweight retrieval pipeline where MRR improves from 0.571 zero-shot SigLIP 2 to 0.760 after text-only LoRA, BM25 fusion and query-bank prior correction on human gold queries.

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar-pro3 is used as the external API model for training-query generation, creating three Korean user-style synthetic queries per product caption for a 7,023-query bank used in LoRA alignment and prior correction.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-declared in a public research-code README; the reported metrics are project results and should be treated as repository-authored claims.

原始记录:K-fashion text-to-image retrieval project uses Solar Pro3 to generate Korean user-style synthetic queries

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zening-cmd 使用 Seedance 2.0 处理多模态内容处理

zening-cmd · Seedance 2.0

A
厂商:ByteDance Seed 模型:Seedance 2.0 来源平台:github 最后复核:2026-06-27T08:22:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zening-cmd 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:22:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:把真实产品 UI 截图、HTML/GSAP 动效、Playwright 渲染帧、Seedance 2.0 数字人和音乐混合为产品 demo、发布或功能介绍 MP4。

公开产物:公开仓库 README 描述了端到端 pipeline:截图采集、HTML 合成、帧渲染、可选 Seedance 2.0 digital human 生成、PIP 合成和音频混音,并列出 Volcengine Ark API key 作为使用 Seedance 2.0 数字人的前置条件。

模型作用:Seedance 2.0 在该管线中生成 AI digital human presenter,用作画中画讲解者,与产品 UI 动画和背景音乐合成为最终产品演示视频。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Seedance 2.0 数字人步骤在 README 中标为 optional,适用于需要数字人讲解的输出;证据为可访问公开仓库自述。

原始记录:zening-cmd 用 Seedance 2.0 数字人制作产品演示视频管线

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CBIhalsen / PolyglotPDF 使用 Grok 2 处理软件工程任务执行

CBIhalsen / PolyglotPDF · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:github 最后复核:2026-06-27T08:22:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CBIhalsen / PolyglotPDF 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:22:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:PolyglotPDF is an open-source PDF/eBook processing application for translating documents while preserving layout, tables, formulas, OCR output, and page structure. Its UI includes a Grok Translate API service and defaul…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public translation tool/repository with demo assets and a configurable Grok backend that can produce translated PDF/eBook output while retaining the original document layout.

模型作用:Grok 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 2 is used as an LLM translation engine: the application sends extracted document text/HTML segments to the selected Grok API backend and uses the returned translation to rebuild layout-preserving PDF/eBook pages.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository configuration and README, not a third-party customer story; it verifies model binding and the public artifact but not production traffic volume.

原始记录:PolyglotPDF uses Grok 2 as a selectable LLM translation backend for layout-preserving PDF/eBook translation

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Josh-XT / AGiXT 使用 Grok Beta 处理软件工程任务执行

Josh-XT / AGiXT · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Josh-XT / AGiXT 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AGiXT implements an xAI provider extension for AI agents, exposing Grok models for text-generation and vision tasks inside the AGiXT provider-rotation system.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public AGiXT repository contains an xAI provider class with friendly_name 'xAI Grok', services ['llm', 'vision'], OpenAI-compatible xAI API URI, and default XAI_AI_MODEL set to 'grok-beta'.

模型作用:Grok Beta 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok Beta is the configured default xAI LLM used by AGiXT agents to generate responses and participate in provider routing when an XAI_API_KEY is supplied.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is code-level integration in a production open-source agent platform, not a customer story; exact model id 'grok-beta' is present in the provider implementation.

原始记录:AGiXT adds xAI Grok Beta as an agent LLM and vision provider

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Skyvern-AI 使用 Seed1.5 VL 处理多模态内容处理

Skyvern-AI · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:23:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Skyvern-AI 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:23:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Skyvern added a UI-TARS client path that calls the Volcano Engine endpoint with model doubao-1-5-thinking-vision-pro-250428 and uses it to generate actions from scraped browser pages for web automation tasks.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The merged Skyvern PR exposes UI-TARS action generation in the Skyvern codebase; the evidence page shows the exact Seed1.5-VL-linked model id and code comments describing generation of actions using UI-TARS (Seed1.5-VL).

模型作用:Seed1.5 VL 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL supplies multimodal screen/page understanding and action selection for GUI/browser automation, turning visual page state into executable Skyvern actions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a GitHub integration PR rather than a customer story; it clearly binds the exact Volcano Engine model id to Seed1.5-VL and a concrete open-source automation artifact.

原始记录:Skyvern integrates Seed1.5-VL / Doubao thinking-vision for UI-TARS browser automation actions

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kaarthik108 / snowChat 使用 Grok 2 处理研究分析和报告生成

kaarthik108 / snowChat · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:github 最后复核:2026-06-27T08:22:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kaarthik108 / snowChat 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:22:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:snowChat is a Streamlit application that lets users ask natural-language questions against Snowflake data. Its model configuration exposes a 'Grok 2' option mapped to grok-2-latest on the xAI API endpoint.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public text-to-SQL application that generates SQL queries and returns Snowflake data/insights through a conversational UI.

模型作用:Grok 2 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Grok 2 acts as one of the LLM backends for interpreting user questions, generating or repairing SQL, and driving the agent workflow that retrieves Snowflake results.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is the public application source and README; the hosted Streamlit demo may redirect, so the repository is used as the stable public artifact.

原始记录:snowChat offers Grok 2 for natural-language querying of Snowflake data

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

langgenius / Dify 使用 Seed1.5 VL 处理软件工程任务执行

langgenius / Dify · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:23:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

langgenius / Dify 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:23:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Dify's official plugin repository added the Doubao-1.5-thinking-vision-pro model to the Volcengine MaaS provider so Dify users can select the ByteDance vision-thinking model inside application workflows.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The merged PR is titled 'feat: add Doubao-1.5-thinking-vision-pro model' and records the model addition in Dify's official plugins artifact for provider-based model invocation.

模型作用:Seed1.5 VL 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL contributes vision-language reasoning capability to Dify apps through the Volcengine MaaS provider, enabling workflows that need image understanding plus text reasoning.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an official plugin integration PR, not an end-customer narrative; it is accepted because it is a public, merged software artifact enabling concrete use of the model in Dify.

原始记录:Dify official Volcengine MaaS plugin adds Doubao-1.5-thinking-vision-pro for multimodal app workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

gusibi / MoliTodo 使用 Grok Beta 处理智能体流程编排

gusibi / MoliTodo · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

gusibi / MoliTodo 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MoliTodo, a desktop productivity/todo application, provides AI settings that let users configure the xAI endpoint and select Grok Beta for in-app AI assistance.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public application code includes an xAI provider settings panel with endpoint placeholder 'https://api.x.ai/v1' and a model selector option value 'grok-beta' labeled 'Grok Beta'.

模型作用:Grok Beta 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Grok Beta supplies the LLM backend option for MoliTodo's AI features once the user configures an xAI API key and chooses the model in settings.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public app implementation and UI configuration rather than a written deployment case study; the exact model id is bound in the settings component.

原始记录:MoliTodo exposes Grok Beta in its built-in AI provider settings

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

bytedance / UI-TARS-desktop and Age… 使用 Seed1.5 VL 处理研究分析和报告生成

bytedance / UI-TARS-desktop and Agent TARS · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:28:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

bytedance / UI-TARS-desktop and Agent TARS 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:28:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Run Agent TARS CLI with provider volcengine and model doubao-1-5-thinking-vision-pro-250428 to execute multimodal agent tasks across terminal, computer, browser and MCP tools, including showcased booking, chart-generati…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public Agent TARS artifact ships CLI/Web UI documentation, showcase videos, and quick-start commands that bind Seed1.5-VL's Volcano Engine model ID to real GUI-agent and browser-operator workflows.

模型作用:Seed1.5 VL 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL contributes multimodal vision and reasoning for GUI/browser-agent perception and planning when users choose the Volcengine provider model ID.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is official open-source product documentation and snapshots that explicitly reference doubao-1-5-thinking-vision-pro-250428; because the project is ByteDance-affiliated, it is less independent than third-party …

原始记录:Agent TARS uses Seed1.5-VL/Doubao Thinking Vision Pro as a Volcengine provider option for multimodal computer and browser agents

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mluggy / Apify Deep Research 使用 Grok 2 处理软件工程任务执行

mluggy / Apify Deep Research · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:github 最后复核:2026-06-27T08:22:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mluggy / Apify Deep Research 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:22:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Apify Deep Research is a public tool that combines Apify data collection with a selectable LLM to refine prompts and generate comprehensive research reports in Markdown or HTML. Its README lists xAI Grok 2.1212 as a sup…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A runnable public research-report generator that outputs Markdown/HTML reports after collecting web data through Apify and prompting the selected LLM.

模型作用:Grok 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 2.1212 is the xAI LLM option responsible for synthesizing crawled material, following the report prompt, and producing the final long-form research report.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository and supported-models documentation; it confirms the exact grok-2-1212 binding and public artifact, but not a named end-customer deployment.

原始记录:Apify Deep Research uses Grok 2.1212 to generate research reports from Apify crawls

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SidhuK / WardenApp 使用 Grok Beta 处理知识检索和问答

SidhuK / WardenApp · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SidhuK / WardenApp 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Warden is a native macOS AI chat client with multi-model support, workspaces, assistants, file/PDF chat, model comparison, search, developer tools, and agents; its configuration includes xAI as a chat-completions provid…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public Warden source defines an xAI provider using 'https://api.x.ai/v1/chat/completions' with defaultModel 'grok-beta' and models ['grok-beta']; the README documents the downloadable native macOS app and its multi-…

模型作用:Grok Beta 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:Grok Beta acts as Warden's configured xAI chat model option, enabling users to route desktop chat and assistant workflows to xAI via their own API key.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence combines README product description with exact provider configuration in code; no separate customer testimonial was found.

原始记录:Warden native macOS AI client supports xAI Grok Beta for chat

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ZhangLab-Kiz / DeepCellSeek 使用 Grok 2 处理研究分析和报告生成

ZhangLab-Kiz / DeepCellSeek · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:github 最后复核:2026-06-27T08:22:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ZhangLab-Kiz / DeepCellSeek 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:22:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DeepCellSeek is an R package and web tool for using LLMs to annotate cell types and subtypes from single-cell RNA-seq marker genes. Its Grok provider configuration lists grok-2-1212 and makes it the default Grok model.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public bioinformatics package/web application that returns LLM-generated cell type annotations and integrates them back into Seurat-style single-cell analysis workflows for visualization.

模型作用:Grok 2 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Grok 2.1212 provides biological reasoning over marker-gene inputs, tissue context, and species metadata to propose cell type/subtype labels as part of the DeepCellSeek annotation pipeline.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence verifies the exact model in the public model configuration and a reachable public web tool; comparative performance claims are not used as the case basis.

原始记录:DeepCellSeek sets Grok 2.1212 as the default Grok model for single-cell RNA-seq annotation

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Hrishikesh332 / Solana-AI-Rewind 使用 Grok Beta 处理研究分析和报告生成

Hrishikesh332 / Solana-AI-Rewind · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Hrishikesh332 / Solana-AI-Rewind 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Solana-AI-Rewind Express backend analyzes Solana wallet data and asks an LLM to produce a concise, humorous Indian-crypto-style roast of the wallet's activity in ten points.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public server code creates a LangChain ChatXAI client with model 'grok-beta' and exposes an '/analyze-wallet' endpoint that returns JSON containing success status, the generated analysis, token usage, and model fing…

模型作用:Grok Beta 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Grok Beta is the generation model invoked on wallet data; it transforms structured wallet statistics into the final natural-language roast and analysis returned by the API.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small project repository with direct executable code evidence; public README is minimal, but the artifact and exact model call are reachable.

原始记录:Solana-AI-Rewind uses Grok Beta to analyze and roast Solana wallet activity

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

EswarDivi / NarrateIt 使用 Grok 2 处理软件工程任务执行

EswarDivi / NarrateIt · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:github 最后复核:2026-06-27T08:22:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

EswarDivi / NarrateIt 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:22:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:NarrateIt is a public Streamlit project that converts text fetched from a URL into an audio file simulating a conversation between two speakers. Its README states that it uses Together AI with the grok-2-1212 model for …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public app/repository that produces a two-speaker narrated audio file from web-page text.

模型作用:Grok 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 2.1212 performs the LLM inference step that turns fetched source text into the dialogue/narration script before Deepgram converts it to speech.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public project README; the repository is used as the stable artifact because the hosted Streamlit app can redirect during automated checks.

原始记录:NarrateIt uses Grok 2.1212 to transform URL text into a two-speaker audio narration

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

browserwing / BrowserWing 使用 Grok Beta 处理软件工程任务执行

browserwing / BrowserWing · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:18:05Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

browserwing / BrowserWing 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:BrowserWing is a browser automation platform that turns websites into structured data through built-in scripts, a CLI for AI agents, MCP/Skills protocol support, and a visual recorder; its agent LLM layer includes an xA…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public BrowserWing repository documents a runnable automation product and its backend agent LLM code maps provider 'xai' to models including 'grok-beta' and 'grok-vision-beta'.

模型作用:Grok Beta 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok Beta is available as an xAI LLM option for BrowserWing's agent layer, contributing language-model reasoning and command generation for browser automation/data-extraction workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is product repository plus backend model mapping; it proves supported use in the agent layer but does not include an external customer story.

原始记录:BrowserWing lists Grok Beta for browser automation agent workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

fal.ai 使用 Seedance 1.0 处理多模态内容处理

fal.ai · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:fal.ai 最后复核:2026-06-27T08:18:05Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

fal.ai 公开的多模态生成与理解案例,来源为 fal.ai,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:fal.ai 将 ByteDance Seedance 1.0 Pro 封装为 Text-to-Video 推理产品,提供网页 Playground、输入表单与 API 调用入口,供开发者用文本提示生成短视频。

公开产物:公开产品页显示该模型为“Seedance 1.0 Pro -- Text to Video”,描述为 ByteDance 开发的高质量视频生成模型,并提供商业使用的推理 API 与示例提示输入。

模型作用:Seedance 1.0 Pro 负责从文本提示合成视频内容;fal.ai 负责托管、API 化、参数表单和开发者调用工作流。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据是模型托管产品页,不是终端客户故事;页面可访问且明确绑定 Seedance 1.0 Pro 与 text-to-video 任务。

原始记录:fal.ai 上架 Seedance 1.0 Pro Text-to-Video 商用推理 API 与 Playground

已有真实案例 多模态生成与理解fal.aiA 类可核验real_case auto_approved 进入模型卡精选

fal.ai 使用 Seedance 1.0 处理多模态内容处理

fal.ai · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:fal.ai 最后复核:2026-06-27T08:18:05Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

fal.ai 公开的多模态生成与理解案例,来源为 fal.ai,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:fal.ai 将 ByteDance Seedance 1.0 Pro 封装为 Image-to-Video API,开发者可上传/提供图像并输入提示词,将静态画面生成动态视频。

公开产物:公开产品页标题为“Seedance 1.0 Pro (Image to Video) API on fal”,正文显示“Seedance 1.0 Pro -- Image to Video”,并提供 Playground、API、输入表单与商业推理入口。

模型作用:Seedance 1.0 Pro 负责保持参考图像主体并生成动态视频镜头;fal.ai 提供模型服务化、调用接口和交互式试用。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据是模型托管产品页,不是第三方客户访谈;页面可访问且明确绑定 Seedance 1.0 Pro 与 image-to-video 任务。

原始记录:fal.ai 上架 Seedance 1.0 Pro Image-to-Video 商用推理 API 与 Playground

已有真实案例 多模态生成与理解fal.aiA 类可核验real_case auto_approved 进入模型卡精选

fal.ai 使用 Seedance 1.0 处理多模态内容处理

fal.ai · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:fal.ai 最后复核:2026-06-27T08:18:05Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

fal.ai 公开的多模态生成与理解案例,来源为 fal.ai,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:fal.ai 将 Seedance 1.0 Lite 的文本生成视频能力做成可调用的推理 API 和在线 Playground,面向需要较轻量视频生成能力的开发者。

公开产物:公开产品页标题为“Seedance 1.0 Lite (Text to Video) API on fal”,正文标注 fal-ai/bytedance/seedance/v1/lite/text-to-video,并提供 API、Playground 和示例提示。

模型作用:Seedance 1.0 Lite 承担从自然语言提示生成视频片段的核心生成步骤;fal.ai 将其部署为低门槛的云端推理服务。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:同属 fal.ai 平台的不同公开产品入口;证据页明确显示 Seedance 1.0 Lite 与 text-to-video,适合作为集成/平台使用案例。

原始记录:fal.ai 提供 Seedance 1.0 Lite Text-to-Video API 给开发者快速生成视频

已有真实案例 多模态生成与理解fal.aiA 类可核验real_case auto_approved 进入模型卡精选

Replicate 使用 Seedance 1.0 处理多模态内容处理

Replicate · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:Replicate 最后复核:2026-06-27T08:18:05Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Replicate 公开的多模态生成与理解案例,来源为 Replicate,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Replicate 将 bytedance/seedance-1-pro 作为公开模型 API 托管,面向开发者提供文本到视频和图像到视频调用能力,可生成 5 秒或 10 秒、480p 或 1080p 的视频。

公开产物:Replicate API 页面显示“ByteDance Seedance 1 Pro | Text to Video API”,描述为支持 text-to-video 和 image-to-video,并显示该模型已有约 2.2M runs。

模型作用:Seedance 1 Pro 负责依据文本或图片输入生成公开视频输出;Replicate 负责 API reference、Playground、示例和模型运行托管。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据是公开 API/产品页,含运行量和任务说明;不是单一品牌客户故事,但属于可访问的真实开发者平台使用。

原始记录:Replicate 托管 bytedance/seedance-1-pro 并提供文本/图像到视频 API

已有真实案例 多模态生成与理解ReplicateA 类可核验real_case auto_approved 进入模型卡精选

Segmind 使用 Seedance 1.0 处理文档理解和结构化处理

Segmind · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:Segmind 最后复核:2026-06-27T08:18:05Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

Segmind 公开的文档理解与处理案例,来源为 Segmind,复核于 2026-06-27T08:18:05Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Segmind 将 Seedance 1.0 Pro 做成 Serverless API,提供 Playground、API 定价与 Python/JavaScript/cURL 示例,供应用把文本和图像转换为动态视频。

公开产物:公开文档页标题为“Seedance 1.0 Pro API Documentation | Segmind”,正文写明“Seedance 1.0 Pro Serverless API”可将文本和图像转为 720p 动态视频,并提供 POST /v2/seedance-pro 调用方式。

模型作用:Seedance 1.0 Pro 负责根据文本/图像输入生成具有电影叙事感的 720p 动态视频;Segmind 提供无服务器 API、SDK 示例和运行入口。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 API 文档/产品页;不是终端用户案例,但页面明确描述 Segmind 对 Seedance 1.0 Pro 的产品化使用和可访问产物。

原始记录:Segmind 提供 Seedance 1.0 Pro Serverless API 用于文本/图像生成动态视频

已有真实案例 文档理解与处理SegmindA 类可核验real_case auto_approved 进入模型卡精选

TheoLeeCJ 使用 Llama 4 Maverick 处理软件工程任务执行

TheoLeeCJ · Llama 4 Maverick

A
厂商:Meta / Llama 模型:Llama 4 Maverick 来源平台:GitHub 最后复核:2026-06-27T08:32:32Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

TheoLeeCJ 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:32:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A human-in-the-loop desktop/browser computer-use agent analyzes screenshots with Meta Llama 4 Maverick and uses UI-TARS for UI element coordinates before proposing actions for the user to accept or reject.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes runnable agent code and a public trajectory explorer showing sample interaction trajectories for the agent.

模型作用:Llama 4 Maverick 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Llama 4 Maverick provides the native vision-language reasoning over screenshots, interpreting UI state and deciding the next action while a separate UI-TARS model supplies click coordinates.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source demo; README warns that actions can be destructive and require close human monitoring. Evidence is from the project README and reachable public trajectory site.

原始记录:TheoLeeCJ built an affordable computer-use agent with Llama 4 Maverick vision

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ThisIs-Developer 使用 Llama 2 Chat 7B 处理代码审查和测试生成

ThisIs-Developer · Llama 2 Chat 7B

A
厂商:Meta / Llama 模型:Llama 2 Chat 7B 来源平台:github 最后复核:2026-06-27T08:30:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ThisIs-Developer 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:30:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:基于《The Gale Encyclopedia of Medicine》PDF 构建医学问答聊天机器人,可回答医学主题问题、总结医学文章并生成医学文本。

公开产物:公开 GitHub 仓库提供 Chainlit/LangChain 医学聊天机器人代码、医学百科 PDF 数据、截图和运行说明。

模型作用:README 明确说明项目使用 Llama-2-7B-Chat-GGML(llama-2-7b-chat.ggmlv3.q2_K.bin)作为生成式聊天模型,结合向量检索回答医学文档问题。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:开源项目案例;README 声明仍在开发且不应替代专业医疗建议;未声称生产部署。

原始记录:ThisIs-Developer 用 Llama-2-7B-Chat-GGML 构建医学 PDF 问答聊天机器人

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LinJT798 使用 Seedance 1.0 处理可玩交互原型构建

LinJT798 · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:github 最后复核:2026-06-27T08:30:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LinJT798 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T08:30:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository implements an AI workflow for game developers: a user enters a character description, chooses generated character art and desired motions such as jump, run, walk, wave, rotate, or idle, and the app genera…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A runnable open-source Python application that produces character animation previews and game-engine-ready sprite sheets; the README shows sample character designs, animation preview GIFs, and sprite-sheet outputs.

模型作用:Seedance 1.0 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:The README identifies the video stage as using a Doubao AI video model / Volcengine Doubao video generation technology, and the GitHub repository description explicitly says it uses doubao-seedance-1-0-pro to produce an…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub repository and README. The repository is a developer-built tool rather than a formal customer story; model binding is strongest in the GitHub repository description naming doubao-seedanc…

原始记录:AI Animation Generator uses Doubao/Seedance to turn character descriptions into game sprite animations

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

talhaanwarch 使用 Llama 2 Chat 7B 处理真实任务执行

talhaanwarch · Llama 2 Chat 7B

A
厂商:Meta / Llama 模型:Llama 2 Chat 7B 来源平台:github 最后复核:2026-06-27T08:30:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

talhaanwarch 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:30:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建可在 CPU-only 低资源 VPS 上运行、具备对话上下文/记忆能力的 Streamlit 聊天应用。

公开产物:公开仓库包含 app.py、Dockerfile、docker-compose.yml、requirements.txt 和部署说明,README 给出项目概览与运行方式。

模型作用:README 明确说明聊天机器人由量化 GGML 版本 Llama-2-7B-Chat 驱动,并通过该模型在资源受限环境中提供文本生成和多轮对话能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 中列出的外部 Working Url 当前 DNS 无法解析;因此 artifact_url 采用可访问的公开 GitHub 仓库。

原始记录:talhaanwarch 用 Llama-2-7B-Chat-GGML 构建带记忆的 Streamlit 聊天机器人

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

NagaYu 使用 Llama 4 Maverick 处理代码审查和测试生成

NagaYu · Llama 4 Maverick

A
厂商:Meta / Llama 模型:Llama 4 Maverick 来源平台:GitHub 最后复核:2026-06-27T08:32:32Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

NagaYu 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:32:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Maverick UI Doctor renders a React component with Playwright, sends both pixels and source code to Llama 4 Maverick, and asks the model to identify visual defects and produce corrected file content.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project exposes a TypeScript CLI/library that returns ranked UI defects and a ready-to-apply unified diff computed locally from the model's corrected file output.

模型作用:Llama 4 Maverick 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Llama 4 Maverick's multimodal reasoning maps visual symptoms, such as invisible CTAs or layout problems, to the relevant source-code causes and proposes code-level fixes.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Requires access to a Llama 4 Maverick deployment through an OpenAI-compatible endpoint; evidence is a public repository README with model ID examples.

原始记录:NagaYu uses Llama 4 Maverick to diagnose and fix React visual regressions

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

WebYeL 使用 Seedance 1.0 处理文档理解和结构化处理

WebYeL · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:github 最后复核:2026-06-27T08:30:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

WebYeL 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T08:30:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The project is a FastAPI plus Vue application for AI image and video generation. Its video module supports text-to-video and image-to-video, multiple resolutions and aspect ratios, adjustable 2-12 second duration, autom…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public full-stack generator application with documented backend APIs such as /api/v1/videos/generate, /api/v1/videos/generate-from-image, task status checks, history listing, deletion, and download endpoints, plus a V…

模型作用:Seedance 1.0 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:The README states the video generation feature is based on doubao-seedance-1-0-pro-fast and lists the video model as doubao-seedance-1-0-pro-fast, making Seedance 1.0 responsible for generating the app's videos from pro…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public repository README with exact model identifier. It is an open-source application artifact, not an independent production customer story; no private deployment metrics are provided.

原始记录:Doubao image/video generator uses doubao-seedance-1-0-pro-fast for text-to-video and image-to-video

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

bestroofingnow 使用 Llama 4 Maverick 处理软件工程任务执行

bestroofingnow · Llama 4 Maverick

A
厂商:Meta / Llama 模型:Llama 4 Maverick 来源平台:GitHub 最后复核:2026-06-27T08:32:32Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

bestroofingnow 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:32:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SiteForge is a conversational service-business website generator that asks about a business, researches the industry, generates marketing content, and creates a complete Next.js website.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a CLI workflow for initializing, building, and previewing generated professional service-business websites, with an estimated total generation cost of about $0.23 per website.

模型作用:Llama 4 Maverick 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Llama 4 Maverick via Groq is used for cost-efficient template expansion and code generation, while Claude handles research, copywriting, and review tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Hybrid multi-model system, so Llama 4 Maverick is one component rather than the only model. README explicitly attributes template expansion and code generation to Llama 4 Maverick via Groq.

原始记录:bestroofingnow's SiteForge uses Llama 4 Maverick via Groq for low-cost website generation

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mowa-ai 使用 Llama 2 Chat 7B 处理真实任务执行

mowa-ai · Llama 2 Chat 7B

A
厂商:Meta / Llama 模型:Llama 2 Chat 7B 来源平台:github 最后复核:2026-06-27T08:30:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mowa-ai 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:30:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:将 Llama-2 7B Chat 部署为可通过 FastAPI 调用的 LLM-as-a-service 服务,并提供接口文档入口。

公开产物:公开仓库提供 laas 服务代码、Poetry 依赖、启动命令、测试说明,并注明在 GCP 单张 Nvidia L4 24GB GPU 上测试。

模型作用:README 明确说明当前版本只支持 7B-chat 模型,并要求下载 llama-2-7b-chat 权重;模型承担服务端对话生成能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:工程服务封装案例;公开证据来自项目 README,未提供外部生产客户或线上服务地址。

原始记录:mowa-ai 将 Llama-2 7B Chat 封装为 FastAPI LLM 服务

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

YiJianRuGu2019 使用 Seedance 1.0 处理浏览器 3D 世界构建

YiJianRuGu2019 · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:github 最后复核:2026-06-27T08:30:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

YiJianRuGu2019 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-27T08:30:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕浏览器 3D 世界构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SAI is a Vue3/TypeScript web application combining AI chat, online search, AI image/video generation, online document generation/editing, and 3D panoramic satellite maps. Its README says the AI image/video function uses…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public front-end application repository that exposes a combined AI productivity interface and includes video-generation capability within the app's feature set.

模型作用:Seedance 1.0 在该案例中承担浏览器 3D 世界构建相关的生成、分析、编排或实现角色。 原始资料写作:Doubao-Seedance-1.0-pro is explicitly named in the README as one of the models used for the app's AI image and video generation feature, contributing the video-generation capability in the application.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub README with exact model name. The README describes the integrated capability but does not include live generated examples or deployment usage metrics.

原始记录:SAI integrates Doubao-Seedance-1.0-pro for AI image and video generation inside a multipurpose web app

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chaoluond 使用 Llama 2 Chat 7B 处理真实任务执行

chaoluond · Llama 2 Chat 7B

A
厂商:Meta / Llama 模型:Llama 2 Chat 7B 来源平台:github 最后复核:2026-06-27T08:30:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chaoluond 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:30:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:训练一个独立模型监控和评估 AI 聊天机器人的回复是否安全,作为类似 OpenAI moderation endpoint 的开源替代方案。

公开产物:公开仓库发布 Safety LLaMA 项目说明、方法论、训练/评估资源和相关代码,用于检测聊天响应中的不安全内容。

模型作用:README 明确写明 Safety LLaMA 是使用 Anthropic red-team harmless 数据集微调得到的 7B-chat LLaMA2 模型,Llama 2 Chat 7B 是安全判别模型的基础。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:属于开源安全研究/工具案例;README 有具体任务和模型绑定,但不是商业客户故事。

原始记录:chaoluond 基于 Llama 2 7B Chat 微调 Safety LLaMA 做聊天安全评估

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

emmytronix 使用 Llama 4 Maverick 处理软件工程任务执行

emmytronix · Llama 4 Maverick

A
厂商:Meta / Llama 模型:Llama 4 Maverick 来源平台:GitHub 最后复核:2026-06-27T08:32:32Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

emmytronix 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:32:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:CodeMentor AI Agent is an interactive programming tutor built for Telex.im with Mastra A2A, offering personalized coding help, explanations, exercises, and gamified learning flows.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains the runnable TypeScript/Express agent, OpenRouter setup, and documentation for deploying the tutor and connecting it to Telex.im.

模型作用:Llama 4 Maverick 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Meta Llama 4 Maverick is the AI model behind the tutor through OpenRouter, generating personalized programming explanations, learning guidance, and conversational responses.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Project documentation says it uses the free OpenRouter tier for Meta Llama 4 Maverick; availability may depend on OpenRouter serving the model.

原始记录:emmytronix built a Telex.im programming tutor with Meta Llama 4 Maverick

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Pikago-hub 使用 Seedance 1.0 处理软件工程任务执行

Pikago-hub · Seedance 1.0

A
厂商:ByteDance Seed 模型:Seedance 1.0 来源平台:github 最后复核:2026-06-27T08:30:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Pikago-hub 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:30:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository provides a GUI application for generating AI videos. Users enter a video prompt, optionally optimize it, choose aspect ratio, resolution, and duration, then click Generate Video while the application show…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Seedance video-generation desktop app that saves generated videos as seedance_video_[timestamp].mp4 and documents typical generation time of 30-60 seconds.

模型作用:Seedance 1.0 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README says the application generates AI-powered videos using the Seedance API from ByteDance and requires an Ark API key from ByteDance; the repository name binds it to seedance1.0pro, so Seedance 1.0 Pro is the vi…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub repository and README. The exact model name is bound primarily by repository name plus Seedance/ByteDance Ark API description; no independent generated-video gallery is included in the R…

原始记录:Seedance Video Generator GUI uses ByteDance Ark Seedance API to create downloadable AI videos

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

growingpenguin 使用 Llama 2 Chat 7B 处理研究分析和报告生成

growingpenguin · Llama 2 Chat 7B

A
厂商:Meta / Llama 模型:Llama 2 Chat 7B 来源平台:github_huggingface 最后复核:2026-06-27T08:30:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

growingpenguin 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:30:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:针对 Naver Sentiment Movie Corpus 的韩文电影评论,将文本判断为正面或负面情感。

公开产物:公开 GitHub 仓库说明训练脚本、数据子集、评测指标和文档;公开 Hugging Face artifact 托管 fine-tuned LLaMA-2-7B Adapter。

模型作用:README 明确说明使用 meta-llama/Llama-2-7b-chat-hf,并通过 LoRA 将其适配为可理解和分析韩文电影评论情感的模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:包含评测指标,但案例核心是公开微调适配器与情感分析任务产物;不是 leaderboard-only。

原始记录:growingpenguin 基于 meta-llama/Llama-2-7b-chat-hf 微调韩文电影评论情感分析适配器

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ihy-adi 使用 Llama 4 Maverick 处理软件工程任务执行

ihy-adi · Llama 4 Maverick

A
厂商:Meta / Llama 模型:Llama 4 Maverick 来源平台:GitHub 最后复核:2026-06-27T08:32:32Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ihy-adi 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:32:32Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Data Analyst Agent processes files such as CSV, Excel, PDF, images, and documents, runs statistical analysis and visualization, and lets users ask natural-language questions about their data.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents a Google Colab-ready data analysis assistant with multi-format ingestion, chart generation, statistical summaries, outlier/correlation detection, and AI-powered Q&A.

模型作用:Llama 4 Maverick 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Together.ai's Llama-4-Maverick-17B-128E-Instruct-FP8 model is used for natural-language understanding and response generation over the computed data-analysis context.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Personal open-source project; deterministic analysis is handled by pandas/matplotlib/seaborn while Llama 4 Maverick handles language understanding and generation.

原始记录:ihy-adi uses Together.ai Llama-4-Maverick-17B for a natural-language data analyst agent

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

University of Washington UW NLP / Q… 使用 Llama 65B 处理真实任务执行

University of Washington UW NLP / QLoRA authors · Llama 65B

A
厂商:Meta / Llama 模型:Llama 65B 来源平台:GitHub + Hugging Face 最后复核:2026-06-27T08:32:44Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

University of Washington UW NLP / QLoRA authors 公开的真实任务执行案例,来源为 公开代码库、Hugging Face 公开空间,复核于 2026-06-27T08:32:44Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:UW NLP 的 QLoRA 项目以 LLaMA 65B 作为底座之一,使用 4-bit quantized pretrained language model 加 LoRA 的方式,在单张 48GB GPU 上微调 65B 参数模型,训练 Guanaco 指令跟随/聊天模型系列。

公开产物:公开产物包括 QLoRA 代码库和 timdettmers/guanaco-65b-merged Hugging Face 模型页面;项目 README 明确描述 Guanaco 作为最佳模型家族,并说明 65B 规模可在 24 小时单 GPU训练完成。

模型作用:LLaMA 65B 提供 65B 参数级别的基础语言能力;QLoRA 在冻结且 4-bit 量化的 LLaMA 65B 上训练 LoRA 适配器,把基础模型转化为可对话的指令模型 Guanaco 65B。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:LLaMA 1 权重许可/访问历史上受限;证据绑定到 65B 参数 LLaMA 和 Guanaco 65B 产物,但部分权重页面可能需要遵守原始 LLaMA 许可。

原始记录:University of Washington 用 LLaMA 65B 通过 QLoRA 训练 Guanaco 65B 指令聊天模型

已有真实案例 真实任务执行公开代码库、Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

alejandrosnz 使用 Grok Beta 处理翻译和本地化处理

alejandrosnz · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:33:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

alejandrosnz 公开的翻译与本地化案例,来源为 公开代码库,复核于 2026-06-27T08:33:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开源命令行工具把 SRT 字幕文件按批次发送给兼容 OpenAI API 的模型翻译成目标语言,并保持原始字幕时间戳;README 给出 xAI Grok 配置:OPENAI_API_URL=https://api.x.ai/v1、OPENAI_MODEL=grok-beta。

公开产物:项目 README 记录了用 xAI grok-beta 测试 1 小时、约 500 句字幕的翻译:初始实现约 6 分钟、成本约 0.25 美元;并给出并发优化后 10/20 线程分别约 50 秒/30 秒的结果。

模型作用:grok-beta 作为字幕翻译生成模型,负责把批量字幕文本翻译成目标语言;工具负责分批、并发、SRT 解析和时间戳保留。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README 的配置与成本/性能实测说明;这是开源工具作者的公开项目,不是官方 xAI 客户案例。

原始记录:SRT LLM Translator 使用 grok-beta 批量翻译字幕并保留时间轴

已有真实案例 翻译与本地化公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DFloat11 / LeanModels 使用 BAGEL 处理真实任务执行

DFloat11 / LeanModels · BAGEL

A
厂商:ByteDance Seed 模型:BAGEL 来源平台:Hugging Face 最后复核:2026-06-27T08:32:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

DFloat11 / LeanModels 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T08:32:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:DFloat11 将 ByteDance-Seed/BAGEL-7B-MoT 作为 base_model,使用 DFloat11 无损压缩方法处理模型权重,并发布可下载的压缩模型与配套使用指南,用于在更低显存占用下运行 BAGEL 的图像生成、图像编辑和多模态能力。

公开产物:公开产物为 DFloat11/BAGEL-7B-MoT-DF11 Hugging Face 模型仓库;模型卡给出 29.21GB 到 19.89GB 的模型大小下降、1024x1024 图像生成峰值显存 30.07GB 到 21.76GB 的对比,并说明输出与原 BFloat16 模型 bit-identical。

模型作用:BAGEL-7B-MoT 提供统一多模态理解与生成基础能力;DFloat11 案例围绕该模型权重做无损压缩,使 BAGEL 能以更低显存部署而保持原始输出一致。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是公开衍生模型/部署优化案例,不是终端客户故事;证据明确声明 base_model 为 ByteDance-Seed/BAGEL-7B-MoT,且产物可访问。

原始记录:DFloat11 发布 BAGEL-7B-MoT 的无损压缩版以降低显存门槛

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

turboderp / ExLlama project 使用 Llama 65B 处理软件工程任务执行

turboderp / ExLlama project · Llama 65B

A
厂商:Meta / Llama 模型:Llama 65B 来源平台:GitHub 最后复核:2026-06-27T08:32:44Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

turboderp / ExLlama project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:32:44Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ExLlama 项目构建了面向量化 LLaMA 模型的本地 GPU 推理运行时,并在 README 中列出 Llama 65B 的双 GPU运行配置和吞吐结果,用于在 4090 + 3090 Ti 等消费/工作站 GPU 上运行 65B 模型。

公开产物:公开 repo 给出可运行代码、配置和 Llama 65B 双 GPU结果表,显示 2,048 token 上下文、约 39.8GB/43.4GB VRAM 配置以及约 16-20 token/s 的生成速度。

模型作用:LLaMA 65B 是该运行时验证和支持的目标大模型之一;ExLlama 的价值在于把 65B 底座通过 GPTQ/量化权重部署到本地多 GPU推理场景。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该案例是开源运行时/部署产物,不是终端业务应用;README 同时包含性能表,但 repo 本身是可访问可运行 artifact,不是单纯 leaderboard。

原始记录:ExLlama 为量化 LLaMA 65B 提供双 GPU 本地推理运行时

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Doriandarko 使用 Grok Beta 处理软件工程任务执行

Doriandarko · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:33:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Doriandarko 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:33:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:o1-engineer 仓库新增 Grok Engineer;grok-eng.py 明确设置 MODEL = "grok-beta",并通过 xAI API 实现 /create、/edit、/add、/review、/planning 等项目文件创建、代码编辑、代码审查和规划命令。

公开产物:产物是可运行的命令行开发助手:根据用户指令生成文件/目录结构、给出编辑指令、应用创建步骤、输出项目计划和代码审查建议。

模型作用:grok-beta 是 Grok Engineer 脚本的核心 chat.completions 模型,负责理解开发者请求、生成代码块、编辑计划、项目规划和审查文本;脚本负责上下文收集、文件写入和交互确认。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自仓库 README 与源码;未验证作者实际调用次数,仅确认公开产物将 grok-beta 绑定为核心模型。

原始记录:o1-engineer 的 Grok Engineer 脚本用 grok-beta 做代码生成、编辑和项目规划

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Intel 使用 BAGEL 处理真实任务执行

Intel · BAGEL

A
厂商:ByteDance Seed 模型:BAGEL 来源平台:Hugging Face 最后复核:2026-06-27T08:32:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Intel 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T08:32:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Intel 使用 intel/auto-round 对 ByteDance-Seed/BAGEL-7B-MoT 生成 mixed int4、group_size 128、symmetric quantization 的量化模型,并在模型卡中给出 vllm-omni 服务方式与 OpenAI-compatible chat completions 调用示例。

公开产物:公开产物为 Intel/BAGEL-7B-MoT-int4-AutoRound 模型仓库;模型卡包含 vllm serve 启动命令,以及生成 sunset.png、将图像转换为 watercolor.png 的 curl 调用示例。

模型作用:BAGEL 作为被量化的统一多模态基础模型,承担文本到图像生成和图像到图像编辑能力;Intel 的 AutoRound 案例展示了对 BAGEL 的低比特部署与 API 化服务路径。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是厂商发布的量化与推理部署产物,不是客户生产故事;证据页面明确写明量化对象是 ByteDance-Seed/BAGEL-7B-MoT,且 artifact 可访问。

原始记录:Intel 用 AutoRound 生成 BAGEL-7B-MoT int4 量化模型并提供 vLLM-Omni 推理示例

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

BigScience Workshop / Petals project 使用 Llama 65B 处理真实任务执行

BigScience Workshop / Petals project · Llama 65B

A
厂商:Meta / Llama 模型:Llama 65B 来源平台:GitHub + Web demo 最后复核:2026-06-27T08:32:44Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

BigScience Workshop / Petals project 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T08:32:44Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Petals 把大模型分片运行在志愿者/组织节点组成的分布式 swarm 中,README 明确列出 Prompt-tune Llama-65B for text semantic classification 的示例,并提供连接 Petals HTTP/WebSocket endpoint 的公开聊天 Web app。

公开产物:公开产物包括 Petals 代码库、LLaMA-65B prompt-tuning Colab 示例、chat.petals.dev Web 应用和 public swarm monitor/source;使用者可以通过 Petals API 进行推理、查看 hidden states 或执行微调。

模型作用:LLaMA 65B 作为 Petals 支持的大参数模型之一,承担文本表示、生成和 prompt-tuning 底座;Petals 通过分布式执行降低单机部署 65B 模型的门槛。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:公开聊天 demo 当前可访问但其在线后端模型可能随时间变化;强证据来自 README 中对 Llama-65B prompt-tuning 示例和 Petals 分布式运行机制的明确描述。

原始记录:BigScience Petals 支持分布式运行并 prompt-tune LLaMA 65B

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zakirkun 使用 Grok Beta 处理文档理解和结构化处理

zakirkun · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:33:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zakirkun 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T08:33:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Deep Eye 是 AI 驱动的渗透测试工具,README 描述其编排 OpenAI、Claude、Grok、Gemini 等模型做智能 payload 生成、45+ 漏洞类型扫描、CVE 情报、误报过滤和合规报告;QUICKSTART 和 config.example.yaml 均给出 Grok 配置 model: "grok-beta"。

公开产物:公开产物支持对目标进行多类漏洞扫描,输出 HTML/PDF/JSON/JUnit/XML/CSV/XLSX 等专业安全报告,并可生成上下文相关 payload、做 AI triage 和 bug bounty 报告。

模型作用:在选择 grok 作为 ai_provider 时,grok-beta 承担安全测试中的自然语言推理与生成任务,包括 payload 生成、扫描结果解释、误报判断和报告内容生成;Deep Eye 框架负责扫描编排、模板、导出和通知。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Grok 是多提供商选项之一,非唯一默认模型;证据明确绑定 grok-beta 配置和工具任务,但不代表所有部署均使用 grok-beta。

原始记录:Deep Eye 渗透测试工具接入 grok-beta 生成安全测试 payload、误报 triage 和报告

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenSenseNova / SenseNova 使用 BAGEL 处理浏览器 3D 世界构建

OpenSenseNova / SenseNova · BAGEL

A
厂商:ByteDance Seed 模型:BAGEL 来源平台:Hugging Face 最后复核:2026-06-27T08:32:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

OpenSenseNova / SenseNova 公开的3D 与 Web 交互案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T08:32:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:SenseNova-SI 项目将 ByteDance-Seed/BAGEL-7B-MoT 作为公开 base_model 之一,使用 SenseNova-SI-8M 空间能力数据体系训练空间智能模型,用于提升多模态模型对视角、位置、3D/空间关系等任务的理解。

公开产物:公开产物为 sensenova/SenseNova-SI-1.1-BAGEL-7B-MoT 模型仓库和 OpenSenseNova/SenseNova-SI 代码仓库;模型卡说明 SenseNova-SI 系列在 VSI-Bench、MMSI、MindCube、ViewSpatial、SITE 等空间智能基准上取得结果,并公开新训练模型以支持研究。

模型作用:BAGEL 的统一理解与生成 MoT 基础能力被用作 SenseNova-SI 家族中的 unified understanding and generation backbone,使项目能在现有多模态基础上扩展空间智能训练。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该案例是研究/开源模型产物;模型卡明确列出 base_model 为 ByteDance-Seed/BAGEL-7B-MoT,任务和公开模型产物可核验。

原始记录:SenseNova-SI 基于 BAGEL 构建空间智能多模态模型

已有真实案例 3D 与 Web 交互Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

rod-trent 使用 Grok Beta 处理文档理解和结构化处理

rod-trent · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub 最后复核:2026-06-27T08:33:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

rod-trent 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T08:33:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:rod-trent 在 JunkDrawer 中发布 Grok Meme Explainer & Infinite Generator:用户上传任意 meme 后,README 明确说明由 grok-beta with vision 返回模板名、文字转写、笑点解释、文化来源和吐槽,再结合本地 Flux/ComfyUI 生成无限变体。

公开产物:README 给出真实输出示例:对 Drake Hotline Bling 梗图返回模板、上下文字转写和对 Rust/Python 梗的解释与吐槽,并说明可基于提示词生成新的 meme 变体。

模型作用:grok-beta 负责视觉/文本理解、meme 模板识别、OCR 式文字转写、笑点解释和生成新变体提示;本地 Flux/ComfyUI 负责图像生成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自作者公开仓库中的子项目说明与示例输出;项目文本将其称为 grok-beta with vision,同时也提到图像生成可切换到 Grok-4,候选仅绑定其 meme 理解/解释环节。

原始记录:Grok Meme Explainer & Infinite Generator 用 grok-beta vision 解释、吐槽并扩展 meme

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Kuleshov Group / MiniLLM project 使用 Llama 65B 处理软件工程任务执行

Kuleshov Group / MiniLLM project · Llama 65B

A
厂商:Meta / Llama 模型:Llama 65B 来源平台:GitHub + Hugging Face 最后复核:2026-06-27T08:32:44Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Kuleshov Group / MiniLLM project 公开的代码代理与软件工程案例,来源为 公开代码库、Hugging Face 公开空间,复核于 2026-06-27T08:32:44Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:MiniLLM 项目提供安装包、命令行工具和预量化模型,用于下载并运行 LLaMA 7B/13B/30B/65B 的 4-bit 权重;README 明确给出 llama-65b-4bit.pt 下载地址和 minillm generate 文本生成流程。

公开产物:公开产物包括 GitHub repo 和 Hugging Face 上的 kuleshov/llama-65b-4bit 模型页;用户可用 MiniLLM 下载 4-bit LLaMA 65B 权重并执行本地文本生成。

模型作用:LLaMA 65B 提供高容量基础语言模型;MiniLLM 的量化和 CLI 封装让 65B 模型能够以 4-bit 形式用于本地生成任务。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是模型运行/量化工具链案例,不是垂直行业业务案例;artifact 公开可访问,且 README 直接列出 llama-65b-4bit。

原始记录:Kuleshov Group MiniLLM 发布 LLaMA 65B 4-bit 权重和命令行生成工具

已有真实案例 代码代理与软件工程公开代码库、Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

OpenSenseNova / SenseNova 使用 BAGEL 处理多模态内容处理

OpenSenseNova / SenseNova · BAGEL

A
厂商:ByteDance Seed 模型:BAGEL 来源平台:Hugging Face / Project Page 最后复核:2026-06-27T08:32:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenSenseNova / SenseNova 公开的多模态生成与理解案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T08:32:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ConsistCompose 在 BAGEL 的 MoT 架构基础上构建统一多模态布局控制框架,把布局坐标嵌入语言提示,实现多实例、身份保持和文本/图像引导的布局可控图像合成。

公开产物:公开产物包括 sensenova/ConsistCompose-BAGEL-7B-MoT 模型仓库、OpenSenseNova/ConsistCompose 代码仓库和项目页面;模型卡说明 ConsistCompose3M 数据集、LELG 方法、Coordinate-CFG,并报告 COCO-Position 上 layout IoU 提升 7.2%、AP 提升 13.7%、Instance Success Ratio 92.6%、Image Success Ratio 76.1%。

模型作用:BAGEL 提供统一理解与生成的 MoT backbone;ConsistCompose 在该基础上加入语言化布局控制训练,使 BAGEL 的图像生成能力扩展到可控多实例组合和身份保持场景。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是公开研究产物和衍生模型案例;证据明确声明 Built upon the MoT architecture of Bagel 且 base_model 为 ByteDance-Seed/BAGEL-7B-MoT。

原始记录:ConsistCompose 基于 BAGEL 做布局可控的多实例图像合成

已有真实案例 多模态生成与理解Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

krumjahn 使用 Grok Beta 处理医疗和生命科学分析

krumjahn · Grok Beta

A
厂商:xAI / Grok 模型:Grok Beta 来源平台:GitHub / PyPI 最后复核:2026-06-27T08:33:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

krumjahn 公开的医疗与生命科学案例,来源为 公开代码库,复核于 2026-06-27T08:33:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:healthai 是 Apple Health AI Analyzer,README 描述其把 Apple Health export.xml 转成可对话的终端健康数据分析助手;CLI 模型清单明确包含 provider_id xai、model xai/grok-beta、label Grok Beta。

公开产物:公开产物可通过 pip/PyPI 安装,读取 Apple Health 数据后让用户用自然语言询问活跃月份、健康趋势、睡眠/运动模式,并生成图表和高保真 CSV/JSON 导出。

模型作用:当用户在 healthai 中选择 xai/grok-beta 时,Grok Beta 负责基于导入的健康数据和工具上下文进行自然语言问答、趋势解释和分析建议;healthai 负责本地数据解析、CLI 会话、图表和导出。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目是多模型工具,grok-beta 是可选模型而非唯一模型;证据明确显示公开模型选项和可安装产物。

原始记录:healthai 将 xai/grok-beta 作为 Apple Health 数据对话分析模型选项

已有真实案例 医疗与生命科学公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Yandex 使用 BAGEL 处理多模态内容处理

Yandex · BAGEL

A
厂商:ByteDance Seed 模型:BAGEL 来源平台:Hugging Face 最后复核:2026-06-27T08:32:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

Yandex 公开的多模态生成与理解案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T08:32:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Yandex 将 ByteDance-Seed/BAGEL-7B-MoT 作为 base_model,在 yandex/alchemist 数据集上进行 T2I fine-tuning,发布 BAGEL-7B-MoT Alchemist,用于生成更高美感和复杂度的图像。

公开产物:公开产物为 yandex/BAGEL-7B-MoT-alchemist 模型仓库;模型卡提供 Alchemist 数据集链接、下载权重代码、加载 BAGEL-Alchemist 的完整示例,并说明该微调版生成 improved aesthetics and complexity 的图像。

模型作用:BAGEL 贡献了统一多模态和文生图基础能力;Yandex 在此基础上用公开 T2I 数据进一步微调,使其成为面向高美感复杂图像生成的 BAGEL 衍生模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是公开模型微调案例,不是第三方客户生产部署;证据页面明确 base_model 为 ByteDance-Seed/BAGEL-7B-MoT,模型和数据集链接可访问。

原始记录:Yandex 用 Alchemist 数据集微调 BAGEL-7B-MoT 提升文生图美感与复杂度

已有真实案例 多模态生成与理解Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

junjunjunbong 使用 Solar Pro 3 处理软件工程任务执行

junjunjunbong · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:github 最后复核:2026-06-27T08:34:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

junjunjunbong 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:34:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Streamlit application for students and tutors that accepts a photo of a wrong math problem, parses the problem, classifies the concept and difficulty, generates five similar problems, writes step-by-step solutions, an…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository documents a working multi-agent pipeline whose final output is an error note containing five similar problems, detailed solutions, formulas, learning advice, and next-study recommendations.

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 3 is explicitly named in the repository as the LLM used by ConceptAgent, ProblemAgent, SolutionAgent, and NoteAgent for concept classification, similar-problem generation, solution writing, and error-note summ…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source project evidence; no independent production usage metrics were found. The README and code bind the artifact to Solar Pro 3, while OCR/extraction stages use other Upstage APIs.

原始记录:SolarNote: Solar Pro 3 multi-agent math error-note generator

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

deep-diver / hf-daily-paper-newslet… 使用 Solar Mini 处理软件工程任务执行

deep-diver / hf-daily-paper-newsletter · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:github 最后复核:2026-06-27T08:37:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

deep-diver / hf-daily-paper-newsletter 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:37:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository implements an automated newsletter for Hugging Face Daily Papers and its Solar integration calls Upstage's Solar endpoint with model="solar-1-mini-chat" to summarize paper titles and abstracts.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public automated newsletter bot/repository that generates daily paper summaries and distributes them through a Google Groups newsletter workflow.

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini is the LLM invoked in the summarization function, converting raw paper titles and abstracts into concise newsletter-ready summaries.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:README text also mentions Gemini for a newer workflow, but the public source file still contains the exact Solar Mini API call; candidate should be treated as code-level evidence for Solar Mini use in this project.

原始记录:Hugging Face Daily Papers newsletter bot uses Solar Mini to summarize research abstracts

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xMetalWorkerx 使用 Grok 2 处理软件工程任务执行

xMetalWorkerx · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:GitHub 最后复核:2026-06-27T08:28:15Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xMetalWorkerx 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:28:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Analyze Google Street View imagery for a user-entered address and classify the property as an abandoned building, vacant lot, valid business, valid residential, or other.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a runnable Python tool that fetches street-level imagery, returns a property category, and includes confidence scores and reasoning for each classification.

模型作用:Grok 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok-2-Vision is the vision analysis component: the README says the tool analyzes images using xAI's Grok-2-Vision model, and the implementation calls model `grok-2-vision-1212` to classify Street View images.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source project evidence; public README and code identify the model and task, but no independent production deployment metrics are provided.

原始记录:Lot Licker uses Grok-2-Vision to classify Google Street View properties

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Samson-U 使用 Solar Pro 3 处理软件工程任务执行

Samson-U · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:github 最后复核:2026-06-27T08:34:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Samson-U 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:34:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A Streamlit tax-filing chatbot that conversationally collects financial inputs, applies compressed tax-code and deduction-guide context, calculates simplified tax liability, and provides personalized tax-saving suggesti…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a runnable Streamlit app with chat logic, API clients, settings, and environment instructions; its README describes the assistant's output as personalized tax preparation guidance and tax-saving …

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README and settings.py bind the OpenRouter model configuration to upstage/solar-pro-3:free by default, making Solar Pro 3 the LLM that interprets user financial inputs and produces conversational tax guidance.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The project is explicitly described as a learning/prototyping demo and not legal or tax advice; public evidence is repo-level rather than a deployed customer story.

原始记录:Tax Preparation Assistant using Solar Pro 3 via OpenRouter

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mahanhrgowda 使用 Grok 2 处理多模态内容处理

mahanhrgowda · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:GitHub 最后复核:2026-06-27T08:28:15Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mahanhrgowda 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:28:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Generate a name-based creative interpretation by mapping English phonemes to chakras, Bhavas, and Rasas, then produce lore text and a matching artistic scene image.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The Streamlit application returns a generated lore narrative for the name and an image prompt/result for a vibrant, detailed artistic scene based on the computed essence.

模型作用:Grok 2 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:The app initializes `xai_text` with model `grok-2-1212` for lore generation and `xai_image` with model `grok-2-image-1212` for image generation, so Grok 2 supplies the user-facing narrative and visual output.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source app evidence from source code and repository description; public deployment status is not independently verified.

原始记录:Name Essence Generator uses Grok-2-1212 for lore and Grok-2 Image for generated scenes

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hunkim / SolarLLMChatDemo 使用 Solar Mini 处理软件工程任务执行

hunkim / SolarLLMChatDemo · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:github 最后复核:2026-06-27T08:37:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hunkim / SolarLLMChatDemo 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:37:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The SolarLLMChatDemo application code uses the OpenAI-compatible Upstage Solar API base URL and calls model="solar-1-mini-chat" for chat completion behavior in the public chat demo project.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public demo repository with linked chat, search, ChatPDF, self-discussion, document-vision, and reasoning Streamlit applications built around Solar LLM usage.

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini provides the conversational generation layer for the demo chat application, producing assistant responses through Upstage's Solar API.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The linked Streamlit app returned an HTTP 303 redirect in command-line verification, so the GitHub repository and exact source file are used as the stable public artifact/evidence.

原始记录:SolarLLM Chat Demo uses Solar Mini for public Streamlit/Gradio chat applications

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AndyNope / Bookitty 使用 Solar Pro 3 处理知识检索和问答

AndyNope / Bookitty · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:github 最后复核:2026-06-27T08:34:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AndyNope / Bookitty 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T08:34:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Bookitty is a Swiss SME accounting web app with Kitty, an AI chatbot that answers accounting, VAT, and app-usage questions, can suggest bookkeeping entries, and can trigger interface actions such as navigation highlight…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public README states that the app is deployed at bookitty.bidebliss.com and that Kitty provides an offline-first knowledge base plus an OpenRouter fallback chain for questions outside the knowledge base; api/chat.ph…

模型作用:Solar Pro 3 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:api/chat.php places upstage/solar-pro-3:free as the first model in the OpenRouter fallback chain, so Solar Pro 3 is the first external LLM attempted for out-of-knowledge-base accounting-chatbot responses before other fr…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Solar Pro 3 is used as first fallback rather than the only model; the artifact is a public product page plus repo evidence, but no per-model live traffic metrics are public.

原始记录:Bookitty Kitty accounting chatbot with Solar Pro 3 first in fallback chain

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DevBauti 使用 Grok 2 处理知识检索和问答

DevBauti · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:GitHub 最后复核:2026-06-27T08:28:15Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DevBauti 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T08:28:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an n8n workflow that downloads a chemistry PDF from Google Drive, chunks and embeds it into Supabase, then answers user questions through a chat interface grounded in the retrieved document context.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public workflow artifact includes an n8n JSON workflow and README describing chat-based Q&A over the processed document, with retrieved context from a Supabase vector store.

模型作用:Grok 2 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly states the chat Q&A uses xAI's `grok-2-1212`; the workflow JSON also contains a model selection for `grok-2-1212`, making Grok 2 the answer-generation model after retrieval.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository evidence is explicit and includes the workflow artifact; no external usage volume or live endpoint is provided.

原始记录:EmbeddingChat uses Grok-2-1212 for document-grounded Q&A in an n8n workflow

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

juyongjiang / CodeSolar 使用 Solar Mini 处理软件工程任务执行

juyongjiang / CodeSolar · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:github 最后复核:2026-06-27T08:37:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

juyongjiang / CodeSolar 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:37:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:CodeSolar documents a workflow to fine-tune base_model_name="solar-1-mini-chat-240612" on the Magicoder-OSS-Instruct-75K coding dataset and publish a codesolar-v1-adapter for solving programming problems.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public project repository containing README instructions, preprocessing, fine-tuning, inference, and adapter upload code for a Solar Mini based coding assistant adapter.

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini is the base chat model being adapted for coding tasks; the project adds code-instruction fine-tuning and uses Solar-as-a-Judge in inference evaluation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small personal research repository with limited stars, but it contains exact model binding and an accessible implementation artifact rather than a generic tutorial page.

原始记录:CodeSolar fine-tunes Solar Mini for coding-problem assistance

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ARUN-L-KUMAR / G-ONE 使用 Solar Pro 3 处理软件工程任务执行

ARUN-L-KUMAR / G-ONE · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:github 最后复核:2026-06-27T08:34:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ARUN-L-KUMAR / G-ONE 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:34:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:G-ONE is a full-stack personal productivity platform that combines calendar, tasks, Gmail, contacts, maps, sheets, and notes context to answer user questions, detect conflicts, prioritize work, and produce optimized dai…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents manual and live planning modes, multi-agent orchestration, persistent AI plans, and a chatbot endpoint that returns replies with context sources, agents used, model source, latency, and fallback…

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:chatbot_router.py exposes Solar Pro 3 as the OpenRouter option openrouter-alt2, routes explicit user selection of that model into the LLM fallback list, maps successful calls back to 'OpenRouter (Solar Pro 3)', and incl…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Solar Pro 3 is one selectable/fallback model in a multi-provider system, not the sole model. Evidence is public source and README documentation, with no public deployment metrics.

原始记录:G-ONE personal task automation chatbot exposes Solar Pro 3 as a selectable OpenRouter model

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Polapo-Invest / Polapo-Invest-PBP 使用 Solar Mini 处理代码审查和测试生成

Polapo-Invest / Polapo-Invest-PBP · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:github 最后复核:2026-06-27T08:37:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Polapo-Invest / Polapo-Invest-PBP 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:37:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Polapo Invest Portfolio Backtesting Platform implements a portfolio analytics web app and configures PredibaseLLM with model_name="solar-1-mini-chat-240612" plus adapter_id="sec-10-k-chatbot" for SEC 10-K risk-facto…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public hackathon project repository for a portfolio backtesting platform that combines asset-allocation analytics with an LLM chatbot for investment insight from SEC filings.

模型作用:Solar Mini 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini is the base model for the fine-tuned SEC 10-K chatbot adapter, enabling natural-language answers over finance disclosures inside the portfolio analysis workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Hackathon/project repository rather than a commercial production deployment, but it includes exact model binding, app code, notebook fine-tuning output, and a public artifact.

原始记录:Polapo Invest uses a Solar Mini fine-tuned SEC chatbot in a portfolio backtesting platform

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

iflow-mcp 使用 Grok 2 处理软件工程任务执行

iflow-mcp · Grok 2

A
厂商:xAI / Grok 模型:Grok 2 来源平台:GitHub 最后复核:2026-06-27T08:28:15Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

iflow-mcp 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:28:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Provide a Model Context Protocol server with a `generate_image` tool so MCP clients can request images from text prompts through the xAI image-generation API.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains a TypeScript MCP server, Docker/deployment instructions, and a generated-image tool that can return generated image URLs or base64 JSON responses.

模型作用:Grok 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The implementation posts to `/images/generations` with model `grok-2-image`, making Grok 2 Image the backend that transforms user prompts into generated images for MCP clients.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public repository and source code verify the model and artifact; operational deployments by downstream MCP users are not measured.

原始记录:GrokArt MCP server exposes Grok-2-image generation as an MCP tool

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

EmilBBbbbbbb 使用 Solar Pro 3 处理软件工程任务执行

EmilBBbbbbbb · Solar Pro 3

A
厂商:Upstage / Solar 模型:Solar Pro 3 来源平台:github 最后复核:2026-06-27T08:34:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

EmilBBbbbbbb 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:34:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A multi-agent technical interview simulator in which interviewer, observer, evaluator, and feedback-generator agents conduct an interview, analyze candidate responses, adapt question difficulty, evaluate correctness, an…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The README describes a runnable CLI workflow that asks for candidate profile details, conducts the interview turn by turn, records internal agent thoughts, supports finish/status commands, and stores interview logs as J…

模型作用:Solar Pro 3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README's environment setup selects LLM_PROVIDER=openrouter and sets OPENROUTER_MODEL=upstage/solar-pro-3:free, binding Solar Pro 3 to the agent LLM calls used for interview question generation, response analysis, ev…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public evidence is from the repository README and appears to be a prototype/demo rather than a production customer deployment; still includes a concrete runnable artifact and exact Solar Pro 3 model binding.

原始记录:Multi-Agent Interview Coach configured with Solar Pro 3

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

infiniflow / RAGFlow 使用 Solar Mini 处理软件工程任务执行

infiniflow / RAGFlow · Solar Mini

A
厂商:Upstage / Solar 模型:Solar Mini 来源平台:github 最后复核:2026-06-27T08:37:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

infiniflow / RAGFlow 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:37:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RAGFlow's public LLM factory configuration includes the Upstage provider and the exact llm_name "solar-1-mini-chat" tagged as an LLM chat model, making Solar Mini selectable in the open-source RAG engine.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A widely used public RAG platform/repository where users can configure Solar Mini as the generation model for document-grounded chat and agentic RAG workflows.

模型作用:Solar Mini 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Mini contributes the chat generation component in RAGFlow deployments that select the Upstage provider, turning retrieved context into end-user answers.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is platform integration evidence rather than a named end-customer story; accepted because it is a public product artifact with exact model binding and a concrete RAG task path.

原始记录:RAGFlow exposes Solar Mini as an Upstage LLM option for RAG applications

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Openmyeyesxz / Image-Renamd_By_AI-O… 使用 Seed1.5 VL 处理软件工程任务执行

Openmyeyesxz / Image-Renamd_By_AI-OCR · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:28:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Openmyeyesxz / Image-Renamd_By_AI-OCR 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:28:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Detect rigid field-trial sample labels with YOLO, crop the label region, send the crop to Ark using model doubao-1-5-thinking-vision-pro-250428, OCR the sample name/code, and batch rename images for phenotypic data coll…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo provides a command-line workflow that writes rename_mapping.csv and moves successfully processed images into renamed output folders using OCR-derived filenames such as sample-code numbered names.

模型作用:Seed1.5 VL 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL, exposed as doubao-1-5-thinking-vision-pro-250428 on Volcano Engine Ark, performs the visual OCR step on cropped label images and returns the text used as the canonical filename base.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub project and code path explicitly sets DEFAULT_ARK_MODEL to doubao-1-5-thinking-vision-pro-250428; usage scale is not independently reported.

原始记录:Image-Renamd_By_AI-OCR uses Seed1.5-VL/Doubao Thinking Vision Pro for field-trial label OCR and batch image renaming

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DSeaStar / ai-diary 使用 Seed1.5 VL 处理研究分析和报告生成

DSeaStar / ai-diary · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:28:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DSeaStar / ai-diary 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:28:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Automatically take hourly desktop screenshots, upload the image, ask doubao-1-5-thinking-vision-pro-250428 to summarize what appears on screen and infer the user's activity, then turn the analysis into diary content.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The application saves generated diary text under a diaries directory and can sync the generated content and screenshot URL into Notion when configured.

模型作用:Seed1.5 VL 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL provides multimodal screenshot understanding: it reads the uploaded screenshot and produces the concise visual activity summary that becomes the input for diary generation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is public source code plus README describing the Doubao screenshot-analysis workflow; repository does not publish production usage metrics.

原始记录:AI Diary uses Seed1.5-VL/Doubao Thinking Vision Pro to analyze screenshots and generate diary entries

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

schxar / LibrarySearch 使用 Seed1.5 VL 处理研究分析和报告生成

schxar / LibrarySearch · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:28:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

schxar / LibrarySearch 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:28:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Provide a Flask VLM service in a broader AI chat, book-search and multimedia-processing platform; the service accepts base64 image data and a prompt, calls Ark with model doubao-1-5-thinking-vision-pro-250428, and retur…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The code returns structured JSON containing the image-analysis text or validation/error messages, enabling the frontend/chat service to handle image inputs alongside text and search workflows.

模型作用:Seed1.5 VL 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL is the visual reasoning backend for uploaded images, converting image_url plus prompt inputs into natural-language analysis results.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence combines README project overview and public VLM service code that hard-codes doubao-1-5-thinking-vision-pro-250428; deployment scale is not stated.

原始记录:LibrarySearch uses Seed1.5-VL/Doubao Thinking Vision Pro for image-content analysis inside an AI chat and library-search platform

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mocehu / image2form_llm 使用 Seed1.5 VL 处理文档理解和结构化处理

mocehu / image2form_llm · Seed1.5 VL

A
厂商:ByteDance Seed 模型:Seed1.5 VL 来源平台:github 最后复核:2026-06-27T08:28:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mocehu / image2form_llm 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T08:28:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Use a multimodal LLM application to convert pictures into form data; the model store exposes doubao-1-5-thinking-vision-pro-250428 under the '豆包多模态深度思考模型' group as a selectable thinking/vision model.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The application artifact is a public Vue/Pinia project whose predefined model list lets users select Seed1.5-VL/Doubao Thinking Vision Pro for image-to-form conversion workflows.

模型作用:Seed1.5 VL 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Seed1.5-VL supplies the multimodal image-understanding and reasoning model option for extracting structured form content from images.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public application repository with an explicit model ID in the product model store and a repo description stating the image-to-form task; runtime adoption is not quantified.

原始记录:image2form_llm offers Seed1.5-VL/Doubao Thinking Vision Pro as a predefined model for converting images into forms

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

fmateoc 使用 Claude 4.1 Opus 处理软件工程任务执行

fmateoc · Claude 4.1 Opus

A
厂商:Anthropic / Claude 模型:Claude 4.1 Opus 来源平台:github 最后复核:2026-06-27T08:49:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

fmateoc 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:49:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an end-to-end Gradle/Java application for a real workplace drudge task: matching messy customer/entity records with rules-based logic derived from the user's domain expertise.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains the generated application, requirements, foundation prompt, Gradle project files, source code and reference data. The author reports that Opus produced substantial code but required manual restru…

模型作用:Claude 4.1 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4.1 was used for the coding/code-generation step after requirements work, producing the initial project implementation and multiple class/file attempts for the entity-matching system.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a first-person README. It is a mixed-success real use case, not a polished customer story; the author explicitly notes manual fixes were needed.

原始记录:fmateoc used Claude Opus 4.1 to generate a rules-based entity matching application

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Cursor 使用 Claude 4 Opus 处理软件工程任务执行

Cursor · Claude 4 Opus

A
厂商:Anthropic / Claude 模型:Claude 4 Opus 来源平台:official_web 最后复核:2026-06-27T08:47:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Cursor 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:47:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Use Claude Opus 4 inside Cursor's developer workflow for coding assistance and complex codebase understanding.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic's launch evidence says Cursor called Opus 4 state-of-the-art for coding and a leap forward in complex codebase understanding; the public Cursor product page is the accessible artifact for the coding tool.

模型作用:Claude 4 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4 contributes frontier coding performance and deeper understanding of large codebases for Cursor's AI-assisted editing and agentic coding workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an Anthropic launch-page customer statement; artifact is the public Cursor product rather than a per-session demo.

原始记录:Cursor uses Claude Opus 4 for complex codebase understanding in its AI code editor

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Cognition 使用 o1-preview 处理软件工程任务执行

Cognition · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:engineering_blog 最后复核:2026-06-27T09:42:50Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Cognition 公开的代码代理与软件工程案例,来源为 博客记录,复核于 2026-06-27T09:42:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Cognition evaluated OpenAI o1-preview in Devin, its autonomous software engineering agent, including tasks that require installing libraries, browsing for context, editing and running code, diagnosing dependency errors,…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Cognition reported that replacing key Devin-Base subsystems with the o1 series produced significant gains on its internal cognition-golden evaluation suite; in a concrete sentiment-analysis task, Devin with o1-preview c…

模型作用:o1-preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Cognition attributes the improvement to o1-preview's ability to reflect, analyze, backtrack, diagnose root causes instead of symptoms, and research online like a human engineer when resolving complex upstream causes in …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The blog is an evaluation and product-engineering writeup rather than a third-party customer deployment; it explicitly names o1-preview and includes concrete Devin task artifacts, but some aggregate benchmark gains are …

原始记录:Cognition tested o1-preview inside Devin for root-cause debugging and autonomous coding-agent tasks

已有真实案例 代码代理与软件工程博客记录A 类可核验real_case auto_approved 进入模型卡精选

ace0109 使用 Claude 4.1 Opus 处理代码审查和测试生成

ace0109 · Claude 4.1 Opus

A
厂商:Anthropic / Claude 模型:Claude 4.1 Opus 来源平台:github 最后复核:2026-06-27T08:49:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ace0109 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T08:49:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create an enterprise NestJS starter scaffold with JWT authentication, RBAC, OAuth, Prisma/PostgreSQL, Redis, WebSocket, scheduled tasks, health checks, Swagger docs, Docker Compose and testing setup.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository and GitHub Pages documentation contain a working TypeScript/NestJS starter template, development instructions, Docker scripts and feature documentation. The repository metadata states it was built …

模型作用:Claude 4.1 Opus 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4.1 was used as the main AI coding assistant to generate and assemble the production-ready backend scaffold and supporting project files.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The explicit model attribution is in the public GitHub repository description/metadata, while the README documents the artifact and technical scope.

原始记录:ace0109 built a production-ready NestJS starter with Claude Opus 4.1

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

cloudstreet-dev 使用 Claude 4.1 Opus 处理软件工程任务执行

cloudstreet-dev · Claude 4.1 Opus

A
厂商:Anthropic / Claude 模型:Claude 4.1 Opus 来源平台:github 最后复核:2026-06-27T08:49:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

cloudstreet-dev 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:49:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Write and structure a beginner-friendly Git and GitHub guide using medical metaphors, covering account creation, Git installation, repositories, cloning, editing, staging, committing and pushing.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains the finished book project 'Git Well Soon' with README, chapter markdown files, practice sections and appendices aimed at Git beginners, students and self-taught developers.

模型作用:Claude 4.1 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4.1 is credited in the public repository description as the model that wrote the book, contributing long-form instructional content and chapter organization.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The model attribution appears in the GitHub repository description; README verifies the public book artifact and its content scope.

原始记录:cloudstreet-dev published a beginner Git book written with Claude Opus 4.1

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Block 使用 Claude 4 Opus 处理软件工程任务执行

Block · Claude 4 Opus

A
厂商:Anthropic / Claude 模型:Claude 4 Opus 来源平台:official_web 最后复核:2026-06-27T08:47:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Block 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:47:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Use Claude Opus 4 in Block's goose coding agent to improve code during editing and debugging.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic's launch evidence says Block found Opus 4 was the first model to boost code quality during editing and debugging in goose while maintaining performance and reliability; goose is publicly accessible as Block's …

模型作用:Claude 4 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4 contributes higher-quality code edits/debugging behavior and reliable agent performance within goose.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an Anthropic launch-page customer statement; artifact is the public goose project site.

原始记录:Block uses Claude Opus 4 in goose for editing and debugging code quality

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

acao 使用 Claude 4.1 Opus 处理软件工程任务执行

acao · Claude 4.1 Opus

A
厂商:Anthropic / Claude 模型:Claude 4.1 Opus 来源平台:github 最后复核:2026-06-27T08:49:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

acao 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:49:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Analyze municipal income taxes, property-tax distributions and economic-development spending patterns in Cuyahoga County, then produce a public research report and interactive web visualization about fiscal flows betwee…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains a substantial markdown report, index.html, conversion scripts and GitHub Actions configuration. The report presents findings such as Cleveland's annual municipal income-tax generation and policy …

模型作用:Claude 4.1 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4.1 is credited in the public repository description as the generation aid for the report/application; the repo's CLAUDE.md describes the report as the source of truth and the generated exports/web app workf…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Model attribution is in repository metadata; artifact files verify the public report and web-app outputs. No independent claim is made about factual correctness of the report analysis.

原始记录:acao generated an interactive Cuyahoga County fiscal-flows report with Claude Opus 4.1

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Augment Code 使用 Claude 4 Opus 处理软件工程任务执行

Augment Code · Claude 4 Opus

A
厂商:Anthropic / Claude 模型:Claude 4 Opus 来源平台:official_web 最后复核:2026-06-27T08:47:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Augment Code 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:47:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Use Claude Opus 4 in Augment Code's AI coding product for complex coding tasks requiring careful, surgical code edits.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic's launch evidence says Augment Code reported higher success rates, more surgical code edits, and more careful work through complex tasks, making Opus 4 the top choice for their primary model.

模型作用:Claude 4 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4 contributes improved success rates, precise code modification, and careful long-running coding behavior for Augment Code's developer assistant workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an Anthropic launch-page customer statement; artifact is Augment Code's public product page.

原始记录:Augment Code uses Claude Opus 4 as a top choice for complex coding tasks

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Cognition 使用 Claude 4 Opus 处理软件工程任务执行

Cognition · Claude 4 Opus

A
厂商:Anthropic / Claude 模型:Claude 4 Opus 来源平台:official_web 最后复核:2026-06-27T08:47:10Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Cognition 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:47:10Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Use Claude Opus 4 in Cognition's Devin software-engineering agent for complex challenges and critical actions in coding workflows.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic's launch evidence says Cognition found Opus 4 excelled at complex challenges other models could not solve and handled critical actions previous models missed; Devin's public product site is reachable as the ar…

模型作用:Claude 4 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4 contributes stronger complex-task resolution and more reliable handling of critical agent actions in Devin's software engineering workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an Anthropic launch-page customer statement; artifact is Devin's public product site.

原始记录:Cognition uses Claude Opus 4 to improve Devin on complex software tasks

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

japatraderdev99 使用 Claude 4.1 Opus 处理软件工程任务执行

japatraderdev99 · Claude 4.1 Opus

A
厂商:Anthropic / Claude 模型:Claude 4.1 Opus 来源平台:github 最后复核:2026-06-27T08:49:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

japatraderdev99 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:49:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Develop a responsive Material Design 3 landing page for F5 Estratégia, a Google Ads agency serving small and medium businesses, with conversion goals, mobile-first layout, GA4/conversion tracking and optimized Web Vital…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository contains the landing-page implementation, index.html variants, Material Web Components styling, analytics/form-validation scripts, Firebase deployment files, image assets and Portuguese README documentati…

模型作用:Claude 4.1 Opus 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Opus 4.1 is credited in the public repository description as the model used to create the landing page, contributing the web implementation and supporting project structure.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The explicit model attribution appears in repository metadata; README and source files verify the public landing-page artifact and task scope.

原始记录:japatraderdev99 built a Google Ads agency landing page with Claude Opus 4.1

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Anthropic 使用 Claude Opus 4.6 (max) 处理软件工程任务执行

Anthropic · Claude Opus 4.6 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.6 (max) 来源平台:github 最后复核:2026-06-27T08:44:38Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Anthropic 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:44:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在人工测试用例指导下,让 Claude Opus 4.6 从零编写一个 Rust 实现的 C 编译器,目标覆盖 x86-64、i686、AArch64 和 RISC-V 64,并尽量不依赖外部编译器工具链。

公开产物:公开仓库包含可构建的 CCC 编译器、前端、代码生成、内置汇编器/链接器、优化与测试说明;README 明确称除一个说明段落外,代码和文档 100% 由 Claude Opus 4.6 编写,并能构建真实 C 代码库。

模型作用:Claude Opus 4.6 负责生成核心代码与文档,实现编译器架构、目标后端、命令行兼容性、测试与构建说明;人类主要提供需通过的测试用例与过程引导。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自 Anthropic 官方 GitHub 仓库 README;模型表述为 Claude Opus 4.6,按任务的 claude-opus-4-6-max 公开模型条目归档。

原始记录:Anthropic 用 Claude Opus 4.6 从零编写多架构 C 编译器

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ThinkOffApp 使用 Claude Opus 4.6 (max) 处理智能体流程编排

ThinkOffApp · Claude Opus 4.6 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.6 (max) 来源平台:github 最后复核:2026-06-27T08:44:38Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ThinkOffApp 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:44:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Wear OS / Galaxy Watch 上构建可语音交互的腕上智能 Agent,把离线语音识别、手表传感器、本地命令路由和云端 Claude Opus 4.6 智能能力结合起来。

公开产物:公开仓库提供 ClawWatch 应用代码、架构说明、演示截图/视频链接、ADB 安装与 API key 配置流程;README 的组件表明确列出 Claude Opus 4.6 作为 cloud intelligence。

模型作用:Claude Opus 4.6 在运行时承担云端智能问答与 Agent 推理部分,处理经 NullClaw-first query path 发送的用户请求,并与本地路由、TTS、传感器数据一起形成手表语音响应。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开 GitHub README;项目同时包含 Vosk、NullClaw、本地命令路由等组件,Claude Opus 4.6 贡献限定为云端智能/推理路径。

原始记录:ThinkOff 在 Galaxy Watch 上构建 ClawWatch 智能手表 Agent

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AGIHunt / Johnixr 使用 Seed-2.1-Pro-Preview 处理软件工程任务执行

AGIHunt / Johnixr · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:web_search 最后复核:2026-06-27T10:23:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AGIHunt / Johnixr 公开的代码代理与软件工程案例,来源为 公开网页,复核于 2026-06-27T10:23:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:AGIHunt 在公开实测文章中将 Seed 2.1 Pro 接入火山方舟 API/Claude Code 类 Agent 工作流,要求模型围绕大学生就业补贴政策从零完成一个公益数据产品:新建公开仓库,检索全国一线、新一线、二线共 49 个城市的就业补贴政策信息,提取图片和文字,整理结构化数据,并用 Vite + React 搭建可搜索、可筛选的网页。

公开产物:产物公开发布为 GitHub 仓库 Johnixr/graduate-subsidy-guide,README 显示项目整理了 242 条高/中可信度政策,覆盖全国 49 个主要城市,提供 CSV、JSON、Markdown 三种数据格式和网页展示;腾讯新闻原文明确给出仓库地址,并描述 Seed 2.1 Pro 完成 383 张图片 OCR、质量审计和前端页面搭建。

模型作用:Seed 2.1 Pro 承担长链路 Agent 与代码工程交付:理解任务目标,规划数据采集流程,检索政府公开信息源,对 383 张图片进行文字识别和字段抽取,生成结构化政策数据,执行质量审计,创建并推送公开 GitHub 仓库,生成 README/.gitignore,并实现 Vite + React 前端及 SVG 图标、多列瀑布流和卡片视觉改进。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自 AGIHunt 在腾讯新闻发布的实测文章和可访问 GitHub 仓库;模型名称在原文中写作 Seed 2.1 Pro/豆包 Seed 2.1 Pro,与本任务 Seed-2.1-Pro-Preview/Pro 家族绑定。案例是作者公开实测与开源产物,不是 benchmark、教程、集合页或发布综述;完整交互日志无法独立审计。

原始记录:AGIHunt 用 Seed 2.1 Pro 从零构建全国大学生就业补贴公益网站

已有真实案例 代码代理与软件工程公开网页A 类可核验real_case auto_approved 进入模型卡精选

bedriyan / Medkit 使用 Claude Opus 4.7 (max) 处理软件工程任务执行

bedriyan / Medkit · Claude Opus 4.7 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.7 (max) 来源平台:GitHub 最后复核:2026-06-27T08:40:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

bedriyan / Medkit 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:40:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A browser-based ER and polyclinic clinical training simulator where learners speak with AI patients, order tests, diagnose, treat, and receive post-session assessment.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository describes an OSCE-style simulator that produces AI patient interactions and an attending-style grading report covering communication, history-taking, and clinical reasoning with guideline citations.

模型作用:Claude Opus 4.7 (max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states that an attending physician powered by Claude Opus 4.7 watches and grades decisions after each session, and that the project was built using Claude Code with Opus 4.7.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the project author's public GitHub README; it is a hackathon/open-source artifact, not an independently audited customer story. Medical cases are described as synthetic and not clinical claims.

原始记录:Medkit uses Claude Opus 4.7 as an attending grader for browser-based clinical training

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Firecrawl 使用 Claude Opus 4.6 (max) 处理软件工程任务执行

Firecrawl · Claude Opus 4.6 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.6 (max) 来源平台:github 最后复核:2026-06-27T08:44:38Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Firecrawl 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:44:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建一个 Chrome 扩展,让用户在结账页右键优惠码输入框后,调用 Firecrawl AI Agent 自动搜索并验证可用优惠码。

公开产物:公开仓库提供扩展源码、安装/构建说明和使用流程;GitHub 仓库描述明确写明该 Chrome extension built with Claude Opus 4.6 Agent Teams and Firecrawl Agents,README 描述了右键搜索、缓存结果、复制/打开来源等可用功能。

模型作用:Claude Opus 4.6 Agent Teams 被用于生成/搭建扩展项目本身,Firecrawl Agents 负责运行时搜索优惠码;最终产物是可本地构建并安装的 Chrome MV3 扩展。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:模型证据位于 GitHub 仓库公开描述,README 主要说明产品功能与 Firecrawl Agent 运行方式;不是 benchmark 或教程。

原始记录:Firecrawl 用 Claude Opus 4.6 Agent Teams 构建优惠码查找 Chrome 扩展

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

adindamochamad / OmniBridge 使用 Claude Opus 4.7 (max) 处理软件工程任务执行

adindamochamad / OmniBridge · Claude Opus 4.7 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.7 (max) 来源平台:GitHub 最后复核:2026-06-27T08:40:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

adindamochamad / OmniBridge 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:40:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:An AI agent for backend/IoT engineers that analyzes legacy serial device traffic, including binary, undocumented, or proprietary protocols, to infer protocol structure and integration logic.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository presents a Tauri/Svelte/Rust application with tests and a public demo video, claiming protocol identification in under a minute for industrial integration workflows.

模型作用:Claude Opus 4.7 (max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly labels the project as built for the Claude Opus 4.7 Hackathon 2026 and displays Claude Opus 4.7 as the model layer for reasoning about raw bytes and protocol behavior.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is self-reported in a public hackathon repository; claims such as speed and protocol coverage should be treated as project claims rather than independently benchmarked results.

原始记录:OmniBridge uses Claude Opus 4.7 to identify undocumented legacy serial-device protocols

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AndyJerry-kk 使用 Seed2.1 Pro 处理智能体流程编排

AndyJerry-kk · Seed2.1 Pro

A
厂商:ByteDance Seed 模型:Seed2.1 Pro 来源平台:github 最后复核:2026-06-27T08:45:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AndyJerry-kk 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:45:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:GitHub 仓库描述明确写明这是一个 Doubao-Seed-2.1-Pro powered AI photography assistant prototype,用于上传照片后获得摄影分析报告和重拍建议。项目目标是验证面向 Sony、Canon、Fujifilm 等实体相机的 AI 拍摄辅助方向。

公开产物:公开仓库 AI-camera-assistant-v0.1 包含可直接打开运行的 index.html、styles.css、app.js 和 README;README 列出照片上传、即时预览、AI 分析、多维评分、改进建议、场景推荐、重拍指导、响应式设计等功能,并提供后续接入真实 VLM API 的接口说明。

模型作用:Seed 2.1 Pro 被用于生成/驱动这个 AI 摄影助手原型的产品与前端实现:把上传图片后的摄影分析需求落成可运行网页 Demo,并组织结构化分析报告、评分维度、建议清单和接口替换方案。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 仓库标题/描述和 README;仓库当前版本包含 mockAnalysis,本身是产品原型和接入架构,未证明已在线调用真实 Seed API 分析用户照片。

原始记录:AndyJerry-kk 用 Doubao-Seed-2.1-Pro 制作 AI 摄影辅助 Web Demo

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

leventilo / Mobius 使用 Claude Opus 4.7 (max) 处理软件工程任务执行

leventilo / Mobius · Claude Opus 4.7 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.7 (max) 来源平台:GitHub 最后复核:2026-06-27T08:40:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

leventilo / Mobius 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:40:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A system where a user drops in a physics paper PDF and receives a Ciechanowski-style interactive HTML simulator based on the paper's physics.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository describes a pipeline that parses the paper, infers physics, composes a typed SimSpec, generates Python solver primitives, runs them in an Anthropic Managed Agents sandbox, and renders an interactive scene…

模型作用:Claude Opus 4.7 (max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states that an Opus 4.7 orchestrator drives a DAG of nine specialist skills and lists the model as claude-opus-4-7 in the stack.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub README for a hackathon/MVP artifact; live demo links are marked pending, so the repository itself is used as the accessible artifact.

原始记录:Mobius uses a Claude Opus 4.7 orchestrator to turn physics papers into interactive simulators

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

taka-avantgarde 使用 Claude Opus 4.6 (max) 处理研究分析和报告生成

taka-avantgarde · Claude Opus 4.6 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.6 (max) 来源平台:github 最后复核:2026-06-27T08:44:38Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

taka-avantgarde 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:44:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建面向风投/并购技术尽调的工具,读取目标代码库,验证创业公司技术主张、识别 AI-washing 风险,并生成结构化技术尽调报告。

公开产物:公开仓库 README 展示 Due Diligence Engine 的安装、使用和输出定位,描述可生成 24 页咨询 PDF;GitHub 仓库描述明确称其为 AI-Powered Technical Due Diligence Engine for Venture Capital,Built on Claude Opus 4.6。

模型作用:Claude Opus 4.6 作为项目构建/推理基础,承担代码库分析、技术风险识别、评分与报告生成流程中的高能力模型角色。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:模型证据来自 GitHub 仓库公开描述,README 也说明可在 Claude Code 等 AI IDE/agent 终端中运行;该案例为公开工具产物。

原始记录:taka-avantgarde 用 Claude Opus 4.6 构建 VC 技术尽调引擎

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

歸藏 / 优设网 使用 Seed2.1 Pro 处理智能体流程编排

歸藏 / 优设网 · Seed2.1 Pro

A
厂商:ByteDance Seed 模型:Seed2.1 Pro 来源平台:web_search 最后复核:2026-06-27T08:45:34Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

歸藏 / 优设网 公开的智能体工作流案例,来源为 公开网页,复核于 2026-06-27T08:45:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:作者通过豆包任务模式和火山引擎 API/Cloud Code Agent 架构实测 Seed 2.1 Pro,要求它在真实内容生产与前端任务中运行复杂 Skills:基于 Seed2.1 官方介绍生成 PPT、生成社交媒体图片卡片,并完成图片 hover 交互、WebGL 贝塞尔曲线、跨整页视差滚动网页等前端动效。

公开产物:公开文章展示并描述 Seed 2.1 Pro 生成的 PPT 页面信息密度、版式和动效表现;社媒卡片可用于封面图/信息卡片/产品介绍图;三个前端任务分别产出横向展开图片交互、带色散的运动贝塞尔曲线、以及用 GSAP、ScrollTrigger 和 Lenis 组织九张图片的连续视差叙事页面。

模型作用:Seed 2.1 Pro 在 Agent 工作流中承担读材料、遵循 Skill 规则、规划页面结构、生成视觉内容和前端代码的核心工作,将抽象视觉/交互描述转化为可运行、可展示的 PPT、图片卡片和网页动效产物。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:文章为作者公开实测和截图展示,产物主要嵌入文章展示,未提供独立 GitHub 仓库或可复现实时代码链接;但不是 benchmark 或教程集合,包含具体使用者、任务和可见产物。

原始记录:设计师歸藏用 Seed 2.1 Pro 跑 PPT Skill、社媒卡片和前端动效 Agent 工作流

已有真实案例 智能体工作流公开网页A 类可核验real_case auto_approved 进入模型卡精选

Ozzy / TUEL AI 使用 Claude Opus 4.6 (max) 处理研究分析和报告生成

Ozzy / TUEL AI · Claude Opus 4.6 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.6 (max) 来源平台:github 最后复核:2026-06-27T08:44:38Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Ozzy / TUEL AI 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T08:44:38Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建一个开源 AI reasoning research platform,把模型推理轨迹保存为可检查、可重跑、可分叉的持久化图,并支持 6-agent swarm、Graph-of-Thoughts 与记忆式工作流。

公开产物:公开仓库提供 Opus Nx 源码、Docker/Node/pnpm 启动方式、UI 截图和功能说明;GitHub 描述明确称 Built with Claude Opus 4.6 by Anthropic,README 展示 persistent reasoning graph、agent swarm、branch comparison 等产物能力。

模型作用:Claude Opus 4.6 被用作构建该研究平台的人机协作/AI research partner,帮助生成和组织多 Agent 推理图平台的代码、架构与文档。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:模型证据来自 GitHub 仓库公开描述及 README 署名区域;README 本身也链接 Anthropic Claude Opus 4.6 新闻页。

原始记录:Ozzy / TUEL AI 用 Claude Opus 4.6 构建持久化推理图平台 Opus Nx

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Saadzwak / Archoff 使用 Claude Opus 4.7 (max) 处理浏览器 3D 世界构建

Saadzwak / Archoff · Claude Opus 4.7 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.7 (max) 来源平台:GitHub 最后复核:2026-06-27T08:40:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Saadzwak / Archoff 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-27T08:40:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕浏览器 3D 世界构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A co-architect product for office interiors that turns a client brief and floor plan into a sourced functional program, 3D test-fit variants, mood board, client argumentaire, slide deck, report, and DXF export.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository describes outputs including costed programs, three 3D SketchUp variants, A3 mood-board PDFs, an 18-slide pitch deck, an A4 report, and a dimensioned A1 DXF with named layers.

模型作用:Claude Opus 4.7 (max) 在该案例中承担浏览器 3D 世界构建相关的生成、分析、编排或实现角色。 原始资料写作:The README states that everything is orchestrated by Claude Opus 4.7, with Vision HD reading plans, managed-agent orchestration producing program/variants/argumentaire, and MCP resources backing decisions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public project README and screenshots in the repo; it is a hackathon/open-source implementation rather than a verified commercial deployment.

原始记录:Archoff uses Claude Opus 4.7 to generate office interior programs, 3D test fits, pitch decks, and DXF exports

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

srichandrak / Maieutic 使用 Claude Opus 4.7 (max) 处理软件工程任务执行

srichandrak / Maieutic · Claude Opus 4.7 (max)

A
厂商:Anthropic / Claude 模型:Claude Opus 4.7 (max) 来源平台:GitHub 最后复核:2026-06-27T08:40:26Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

srichandrak / Maieutic 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:40:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A pedagogical IDE for programming education that asks students to specify intended behavior before coding, supports reasoning-oriented coding conversations, compares submitted code against the student's specification, a…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository describes a Next.js classroom app with student and teacher flows, live dashboard summaries, specification clarification, code/spec comparison, and teacher-facing class pattern detection.

模型作用:Claude Opus 4.7 (max) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README lists @anthropic-ai/sdk with model claude-opus-4-7 and states that Opus is called at seven moments across student and teacher experiences, including spec clarification, coding guidance, code/spec comparison, …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the author's public GitHub README; it is an MVP-style educational tool and not a third-party deployment report.

原始记录:Maieutic uses Claude Opus 4.7 to coach programming students and summarize learning signals for teachers

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Lovable 使用 Claude Opus 4.5 处理真实任务执行

Lovable · Claude Opus 4.5

A
厂商:Anthropic / Claude 模型:Claude Opus 4.5 来源平台:official_web 最后复核:2026-06-27T08:49:08Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Lovable 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T08:49:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Lovable 将 Claude Opus 4.5 提供给用户在 Chat mode 中进行项目规划、需求澄清、迭代设计,并把更好的规划结果衔接到后续代码生成流程。

公开产物:Lovable 官方博客发布“Lovable Now Supports Claude Opus 4.5”,Anthropic 发布页中 Lovable CTO Fabian Hedin 说明 Opus 4.5 在 Lovable chat mode 中提供 frontier reasoning,帮助用户规划和迭代项目。

模型作用:Claude Opus 4.5 负责长链路推理和项目规划,把用户意图转化为更清晰的实施方案,从而提升 Lovable 代码生成前的计划质量。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自 Lovable 官方发布页和 Anthropic 发布页的客户引语;未公开单个终端用户项目链接,但 Lovable 产品页为可访问产物。

原始记录:Lovable 在 Chat mode 中接入 Claude Opus 4.5 做应用规划与迭代

已有真实案例 真实任务执行官方页面A 类可核验real_case auto_approved 进入模型卡精选

Warp 使用 Claude Opus 4.5 处理软件工程任务执行

Warp · Claude Opus 4.5

A
厂商:Anthropic / Claude 模型:Claude Opus 4.5 来源平台:official_web 最后复核:2026-06-27T08:49:08Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Warp 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:49:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Warp 为用户在终端 AI 工作流中接入 Claude Opus 4.5,重点用于 /plan、Planning Mode 等需要持续推理、多步执行和长周期自主处理的开发任务。

公开产物:Warp 官方博客宣布 Claude Opus 4.5 面向所有 Warp 用户可用,并说明它适合 long-horizon tasks;Anthropic 发布页中 Warp 创始人 Zach Lloyd 称其在复杂工作流中减少 dead-ends,并在 Terminal Bench 上较 Sonnet 4.5 提升 15%。

模型作用:Claude Opus 4.5 作为 Warp AI 的规划和推理模型,为终端中的复杂多步任务生成计划、维持上下文并辅助执行。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:Warp 博客和 Anthropic 客户引语均可核验;量化提升来自 Warp/Anthropic 公布的评估口径,非独立第三方复测。

原始记录:Warp 在 /plan 与 Planning Mode 中使用 Claude Opus 4.5 处理长周期开发任务

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

GitHub Copilot 使用 Claude Opus 4.5 处理软件工程任务执行

GitHub Copilot · Claude Opus 4.5

A
厂商:Anthropic / Claude 模型:Claude Opus 4.5 来源平台:official_web 最后复核:2026-06-27T08:49:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

GitHub Copilot 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T08:49:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:GitHub Copilot 早期测试 Claude Opus 4.5,用于重型 agentic coding workflows,尤其是代码迁移和代码重构任务。

公开产物:Anthropic 发布页引用 GitHub Chief Product Officer Mario Rodriguez:Claude Opus 4.5 powers heavy-duty agentic workflows with GitHub Copilot,在内部编码基准上超过既有结果并将 token 使用量减半,尤其适合 code migration 和 code refactoring。

模型作用:Claude Opus 4.5 为 Copilot 的多步编码代理提供高质量代码生成、规划和工具调用能力,帮助迁移和重构任务以更少 token 完成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:原始证据是 Anthropic 发布页中的 GitHub 高管客户引语;artifact_url 为公开 GitHub Copilot 产品页,未提供具体公开 PR 链接。

原始记录:GitHub Copilot 使用 Claude Opus 4.5 支撑代码迁移与重构类 agentic workflow

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Cursor 使用 Claude Opus 4.5 处理智能体流程编排

Cursor · Claude Opus 4.5

A
厂商:Anthropic / Claude 模型:Claude Opus 4.5 来源平台:official_web 最后复核:2026-06-27T08:49:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Cursor 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T08:49:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Cursor 将 Claude Opus 4.5 用作编辑器内编码模型,面向困难 coding tasks 提供更高智能和更优价格表现。

公开产物:Anthropic 发布页引用 Cursor CEO Michael Truell:Claude Opus 4.5 is a notable improvement over prior Claude models inside Cursor,在困难编码任务上具备更好的 pricing and intelligence;Cursor 文档站也列出 claude-opus-4-5 模型条目。

模型作用:Claude Opus 4.5 在 Cursor 中承担代码理解、生成、修改和复杂任务推理,提升困难编码任务处理能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据主要来自 Anthropic 发布页客户引语和 Cursor 可访问产品/文档;未公开单个用户项目产物。

原始记录:Cursor 在困难编码任务中集成 Claude Opus 4.5

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

Notion 使用 Claude Opus 4.5 处理智能体流程编排

Notion · Claude Opus 4.5

A
厂商:Anthropic / Claude 模型:Claude Opus 4.5 来源平台:official_web 最后复核:2026-06-27T08:49:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Notion 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T08:49:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Notion 将 Opus 4.5 提供给 Notion Agent,用于理解用户真实意图并生成可分享的文档、计划或工作内容。

公开产物:Anthropic 发布页引用 Notion AI Lead Engineer Sarah Sachs:Opus 4.5 excels at interpreting what users actually want, producing shareable content on the first try;这是 Notion 首次在 Notion Agent 中提供 Opus。

模型作用:Claude Opus 4.5 负责意图理解、上下文推理和内容生成,帮助 Notion Agent 更快产出可直接分享的工作成果。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:原始证据为 Anthropic 发布页客户引语;artifact_url 为 Notion AI 产品页,具体用户产出未公开。

原始记录:Notion Agent 使用 Claude Opus 4.5 一次生成可分享内容

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

Max Harar 使用 Claude 4.5 Sonnet 处理软件工程任务执行

Max Harar · Claude 4.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 4.5 Sonnet 来源平台:github 最后复核:2026-06-27T08:54:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Max Harar 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:54:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an interactive airline customer-service agent against Sierra's open tau-bench airline domain, using the policy document as the system prompt, seeded airline data as state, and typed tools for tasks such as booking…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes a working single-page live demo at maxharar.com/sierra and documents five scenario chips plus a Pass^1 measurement for claude-sonnet-4-5 on the selected airline tasks.

模型作用:Claude 4.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Sonnet 4.5 is listed as the reasoning model in the stack and is described as driving the agent's policy-aware decisions and tool use over the airline-domain workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The domain is based on an open benchmark environment, but the submitted artifact is a concrete public demo/repository rather than a leaderboard-only entry; repo and live demo were reachable during collection.

原始记录:Max Harar built a live Sierra tau-bench airline customer-service agent with Claude Sonnet 4.5

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Mohamed Sorour 使用 Claude 4.5 Sonnet 处理软件工程任务执行

Mohamed Sorour · Claude 4.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 4.5 Sonnet 来源平台:github 最后复核:2026-06-27T08:54:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Mohamed Sorour 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:54:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a natural-language AWS operations assistant that monitors and optimizes cloud resources through Amazon Bedrock AgentCore Gateway, exposing CloudWatch, Logs, and EBS APIs as tools instead of requiring manual dashb…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains a runnable agent, screenshots, setup instructions, local CLI flow, semantic tool discovery, and documented access to 137 AWS tools for monitoring metrics, analyzing logs, and optimizing EB…

模型作用:Claude 4.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README names Claude Sonnet 4.5 as the LLM responsible for natural-language understanding inside the AgentCore Gateway plus Strands Agents architecture.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The README describes the project as production-ready while AgentCore Runtime deployment is still marked in progress; accepted as a public implementation artifact with clear model binding and task evidence.

原始记录:Mohamed Sorour built an AWS Resource Optimizer Agent with Claude Sonnet 4.5 and Bedrock AgentCore

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

asimniazi63 使用 Claude 4.5 Sonnet 处理软件工程任务执行

asimniazi63 · Claude 4.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 4.5 Sonnet 来源平台:github 最后复核:2026-06-27T08:54:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

asimniazi63 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:54:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an autonomous Enhanced Due Diligence research system that gathers intelligence on people and entities, analyzes risks, maps relationships, and generates comprehensive investigation outputs through a LangGraph-base…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes the DeepAgents codebase and architecture for a production-ready due-diligence research agent, including orchestrator, Claude service, search service, result aggregation, error handling, and repo…

模型作用:Claude 4.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly lists Claude Sonnet 4.5 in the multi-model architecture and assigns it to strategic analysis, query generation, and reflection steps in the due-diligence workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The system also uses GPT-4o, so the case is multi-model; the Claude Sonnet 4.5 contribution is separately identified in the README and tied to specific workflow responsibilities.

原始记录:DeepAgents uses Claude Sonnet 4.5 for autonomous enhanced due-diligence investigations

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Mark2009 使用 Claude 4.5 Sonnet 处理软件工程任务执行

Mark2009 · Claude 4.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 4.5 Sonnet 来源平台:github 最后复核:2026-06-27T08:54:40Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Mark2009 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:54:40Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Port the jsonrepair JavaScript library to C++ and add repair capabilities for malformed JSON documents, producing a usable C++ implementation with examples.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains the C++ jsonrepaircpp library, feature list, examples showing invalid JSON repaired into valid JSON, and attribution that the port was produced by Claude Sonnet 4.5 with additional repair …

模型作用:Claude 4.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README states that the C++ port was created by Claude Sonnet 4.5, binding the model to the software translation and implementation work.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is the project README and public repository; no independent production deployment is claimed, but the artifact is a reachable code library rather than a tutorial, benchmark, or collection page.

原始记录:Mark2009 used Claude Sonnet 4.5 to port jsonrepair from JavaScript to C++

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

KamakuraCrypto 使用 GPT-5.4 (xhigh) 处理代码审查和测试生成

KamakuraCrypto · GPT-5.4 (xhigh)

A
厂商:OpenAI 模型:GPT-5.4 (xhigh) 来源平台:github 最后复核:2026-06-27T09:20:44Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

KamakuraCrypto 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T09:20:44Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Claude Code skills that send a plan Claude just wrote to OpenAI's codex exec CLI for a structured premortem, using a dual-model setup that explicitly includes gpt-5.4 at xhigh reasoning alongside gpt-5.5.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides installable /codex-review and /codex-review-full Claude Code skills, sandbox/full-access modes, convergence-loop instructions, blocker/question output schemas, and setup notes for reviewin…

模型作用:GPT-5.4 (xhigh) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.4 at xhigh reasoning is one of the two Codex reviewers in the premortem loop; it independently examines the proposed plan, reports blockers, likely failure modes, questions, confidence, and shippability status so …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub repository and README; it demonstrates a concrete reusable developer-tool artifact, though production usage outside the repository is not independently verified.

原始记录:KamakuraCrypto uses GPT-5.4 xhigh in a dual-Codex premortem review skill for Claude Code plans

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

NVIDIA 使用 Qwen3.5 397B A17B 处理知识检索和问答

NVIDIA · Qwen3.5 397B A17B

A
厂商:Qwen / Alibaba 模型:Qwen3.5 397B A17B 来源平台:NVIDIA Build / NIM 最后复核:2026-06-27T08:59:46Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

NVIDIA 公开的知识库与检索问答案例,来源为 NVIDIA Build / NIM,复核于 2026-06-27T08:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:NVIDIA 将 Qwen/Qwen3.5-397B-A17B 接入 NVIDIA NIM/Build,提供 /chat/completions 形式的多模态推理 API、Playground 与部署说明,供开发者调用模型处理视觉语言理解、聊天、RAG 和 agentic 任务。

公开产物:公开产品页显示 qwen3.5-397b-a17b Model by Qwen | NVIDIA NIM,并给出 integrate.api.nvidia.com/v1 端点、OpenAI-compatible 工具定义、LangChain ChatNVIDIA 示例和 Linux Docker 部署资料。

模型作用:Qwen3.5 397B A17B 是该 NIM artifact 的核心推理模型,承担文本/图像/视频输入理解与聊天补全输出;NVIDIA 主要提供托管 API、NIM 包装、示例代码和部署运行时。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据是公开产品/API 页面,属于模型托管与部署用例;不是终端客户故事,但有可访问 artifact 和明确组织。

原始记录:NVIDIA NIM 上线 Qwen3.5-397B-A17B 多模态推理 API

已有真实案例 知识库与检索问答NVIDIA Build / NIMA 类可核验real_case auto_approved 进入模型卡精选

devdotbo 使用 GPT-5.4 (xhigh) 处理软件工程任务执行

devdotbo · GPT-5.4 (xhigh)

A
厂商:OpenAI 模型:GPT-5.4 (xhigh) 来源平台:github 最后复核:2026-06-27T08:59:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

devdotbo 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:59:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Claude Code plugin that delegates investigation, bug fixing, feature implementation, refactoring, read-only reviews, adversarial reviews, and dual Codex/Claude reviews to Codex jobs.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public fork exposes slash commands such as /codex:execute, /codex:status, /codex:result, /codex:adversarial-review, and /codex:dual-review, with background job handling and stored final outputs.

模型作用:GPT-5.4 (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The plugin states execute tasks run in YOLO full-access mode with GPT 5.4 at xhigh reasoning; the common configuration section says the fork forces GPT 5.4 at xhigh reasoning for all execute tasks and uses it as the Cod…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository README and repo page; deployment usage beyond the repository is not independently audited.

原始记录:devdotbo forces GPT 5.4 xhigh in a Claude Code Codex YOLO plugin for execution and dual review

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SiliconFlow 使用 Qwen3.5 397B A17B 处理真实任务执行

SiliconFlow · Qwen3.5 397B A17B

A
厂商:Qwen / Alibaba 模型:Qwen3.5 397B A17B 来源平台:SiliconFlow 最后复核:2026-06-27T08:59:46Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

SiliconFlow 公开的真实任务执行案例,来源为 SiliconFlow,复核于 2026-06-27T08:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:SiliconFlow 在模型市场中上线 Qwen/Qwen3.5-397B-A17B,提供 Open in Playground、API Reference、Serverless 计费与调用入口,用于长上下文、多模态、工具调用和推理型应用。

公开产物:公开页面列出 Qwen3.5-397B-A17B、Qwen/Qwen3.5-397B-A17B、Available Serverless、Input/Output token 价格、Playground 和 API Usage,并描述其 397B total/17B activated、256K context、201 languages 等能力。

模型作用:Qwen3.5 397B A17B 为 SiliconFlow 该 serverless API 的实际后端模型,负责处理开发者请求中的视觉语言理解、长上下文分析、工具调用与 reasoning 输出。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开模型产品页;页面含 benchmark 区块,但本候选只使用其 API/Playground/serverless artifact 信息,不把 benchmark 当案例。

原始记录:SiliconFlow 将 Qwen3.5-397B-A17B 作为 Serverless 模型 API 提供

已有真实案例 真实任务执行SiliconFlowA 类可核验real_case auto_approved 进入模型卡精选

Christopher-Schulze / ClankWork 使用 Qwen3.6 Plus 处理智能体流程编排

Christopher-Schulze / ClankWork · Qwen3.6 Plus

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Plus 来源平台:github 最后复核:2026-06-27T08:58:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Christopher-Schulze / ClankWork 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T08:58:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:将 Qwen Code CLI 包装成常驻 macOS 后台 daemon,通过 Telegram 接收任务,并使用 MCP、浏览器自动化、桌面工作流和文档处理能力执行个人工作站自动化任务。

公开产物:公开 GitHub 仓库提供了 ClankWork 原型、安装与配置说明,README 明确说明推荐模型路径使用 Qwen3.6 Plus,并说明其用于长上下文、多模态和函数调用工作流。

模型作用:Qwen3.6 Plus 作为模型层,为本地工作站智能体提供 1M 上下文、多模态理解、函数调用和低成本推理路线,支撑 Telegram 任务入口、视觉/文档工作流和长期上下文保持。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README 的自述,未看到第三方生产部署证明;但仓库和说明公开可访问,模型、使用者、任务与产物均明确。

原始记录:ClankWork 用 Qwen3.6 Plus 构建常驻 macOS Telegram 工作站智能体

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

noosi21 使用 GPT-5.4 (xhigh) 处理软件工程任务执行

noosi21 · GPT-5.4 (xhigh)

A
厂商:OpenAI 模型:GPT-5.4 (xhigh) 来源平台:github 最后复核:2026-06-27T08:59:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

noosi21 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:59:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Smart-contract audit CLI and Codex workstation for Solidity and DeFi protocols, combining skill-based audit prompts, parsing, symbolic/static tools, fuzzing, invariant checks, GitHub URL targets, and Foundry proof-of-co…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository provides npm/Docker/Kali setup, codex-sol audit commands, 35 audit skills, GitHub target auditing, CI/diff audit options, and examples for checking reentrancy, flash loans, upgradability, and bug-bount…

模型作用:GPT-5.4 (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:README binds the LLM path to GPT-5.4 xhigh: commands include --reasoning-effort xhigh, the Codex CLI integration is labeled GPT-5.4 xhigh Agent, and the agent reads AGENTS.md plus skill definitions before executing the …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public tool repository; specific vulnerability findings against third-party targets are examples rather than verified production audit reports.

原始记录:noosi21 uses GPT-5.4 xhigh in Codex Solidity for smart-contract audit workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenRouter 使用 Qwen3.5 397B A17B 处理智能体流程编排

OpenRouter · Qwen3.5 397B A17B

A
厂商:Qwen / Alibaba 模型:Qwen3.5 397B A17B 来源平台:OpenRouter 最后复核:2026-06-27T08:59:46Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

OpenRouter 公开的智能体工作流案例,来源为 OpenRouter,复核于 2026-06-27T08:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:OpenRouter 将 Qwen3.5 397B A17B 纳入模型目录,提供兼容 OpenAI SDK 的统一 API、模型 slug、价格/用量信息和多供应商 endpoint 路由,让开发者通过 OpenRouter 调用该模型完成聊天、代码、agent、图像/视频理解等任务。

公开产物:公开页面标题为 Qwen3.5 397B A17B - API Pricing & Benchmarks | OpenRouter,页面说明 OpenRouter API is OpenAI-compatible,并在页面数据中列出 model slug qwen/qwen3.5-397b-a17b、hf_slug Qwen/Qwen3.5-397B-A17B 以及 DigitalOcean endpoint。

模型作用:Qwen3.5 397B A17B 是 OpenRouter 暴露给开发者的可调用模型;OpenRouter 贡献统一路由、兼容 API、价格展示和 provider endpoint,模型本身贡献语言理解、逻辑推理、代码生成、agent 与视觉理解能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是模型 API 聚合/路由产品中的真实上线 artifact;不是客户应用案例,但组织、任务、公开 artifact 与 exact model slug 均明确。

原始记录:OpenRouter 为 Qwen3.5 397B A17B 提供 OpenAI-compatible 路由 API

已有真实案例 智能体工作流OpenRouterA 类可核验real_case auto_approved 进入模型卡精选

LeopardCode.AI 使用 Qwen3.6 Plus 处理浏览器 3D 世界构建

LeopardCode.AI · Qwen3.6 Plus

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Plus 来源平台:github 最后复核:2026-06-27T08:58:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LeopardCode.AI 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-27T08:58:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发并维护一个浏览器端 3D 空间音频体验构建器,支持在 2D/3D 画布放置声源、HRTF 空间化、姿态物理、关键帧动画、导出与分享。

公开产物:项目 README 提供公开仓库与线上 Vercel demo,说明该应用由 LeopardCode.AI 维护,是 fully agentic AI development 实验,并明确标注 vibecoded with Qwen3.6 Plus。

模型作用:Qwen3.6 Plus 用于代理式/代码生成式开发该 Web 应用,帮助实现交互式空间音频编辑、测试与部署相关功能。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:模型贡献为项目 README 自述;线上 demo 和 GitHub 仓库均可访问,适合作为公开产物型案例。

原始记录:LeopardCode.AI 用 Qwen3.6 Plus 代理式开发 3D Sound Journey Builder

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

NVIDIA 使用 Qwen3.5 397B A17B 处理软件工程任务执行

NVIDIA · Qwen3.5 397B A17B

A
厂商:Qwen / Alibaba 模型:Qwen3.5 397B A17B 来源平台:Hugging Face 最后复核:2026-06-27T08:59:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

NVIDIA 公开的代码代理与软件工程案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T08:59:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:NVIDIA 使用 Model Optimizer 将 Alibaba Qwen/Qwen3.5-397B-A17B 制作为 NVFP4 量化版本,面向 AI Agent、chatbot、RAG 和其他应用部署场景提供预量化模型 artifact。

公开产物:Hugging Face repo nvidia/Qwen3.5-397B-A17B-NVFP4 公开可访问;API 元数据包含 base_model:Qwen/Qwen3.5-397B-A17B、base_model:quantized:Qwen/Qwen3.5-397B-A17B、ModelOpt、FP4 标签;README 写明该模型是 Alibaba Qwen3.5-397B-A17B 的 quantized version,支持 vLLM/SGLang 与 NVIDIA Blackwell。

模型作用:Qwen3.5 397B A17B 提供原始 397B/17B MoE 能力和文本/图像/视频输入能力;NVIDIA 的贡献是用 Model Optimizer 量化并打包为更适合 GPU 部署的 NVFP4 artifact。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是公开可下载的工程 artifact 与部署用例,非 benchmark;与原模型绑定通过 Hugging Face base_model 标签和 README 明确确认。

原始记录:NVIDIA 发布 Qwen3.5-397B-A17B 的 NVFP4 量化部署 artifact

已有真实案例 代码代理与软件工程Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

rwightman 使用 GPT-5.4 (xhigh) 处理软件工程任务执行

rwightman · GPT-5.4 (xhigh)

A
厂商:OpenAI 模型:GPT-5.4 (xhigh) 来源平台:github 最后复核:2026-06-27T08:59:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

rwightman 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:59:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a clean standalone PyTorch implementation of the Gemma 4 model family, including text, vision, and audio towers, generation, KV-cache, multimodal prompt expansion, and checkpoint conversion from Hugging Face safe…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository contains the gemma4-pytorch-codex package with working text generation, image preprocessing and image-conditioned generation, audio preprocessing, save/load, conversion utilities, tests, and documentat…

模型作用:GPT-5.4 (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The GitHub repository title/description states the implementation was built using Codex with GPT-5.4 xhigh, tying the model to the code-generation and implementation task that produced the public artifact.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The exact GPT-5.4 xhigh evidence appears in the GitHub repo title/description rather than the README body; artifact contents are publicly accessible.

原始记录:rwightman produced a standalone Gemma 4 PyTorch implementation using Codex with GPT-5.4 xhigh

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zjcolacode 使用 Qwen3.6 Plus 处理多模态内容处理

zjcolacode · Qwen3.6 Plus

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Plus 来源平台:github 最后复核:2026-06-27T08:58:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zjcolacode 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:58:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:输入任意视频后,系统识别视频类型、按语义切片并评分,依据目标时长筛选片段、去重并重排,再用 ffmpeg 输出精华短视频和结构化 JSON。

公开产物:公开仓库 README 明确说明该 MVP 基于阿里云百炼 Coding Plan 与 qwen3.6-plus 视觉模型;架构中 VisionAnalyst 将抽帧后的多帧图片喂给 qwen3.6-plus,输出结构化片段标注,最终生成 highlight.mp4 和 highlight.json。

模型作用:Qwen3.6 Plus 负责多帧视觉理解、视频类型识别、片段语义标注和信息密度/情绪强度/独立可观看性评分,是自动剪辑决策链路的核心模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目为 MVP,未声明大规模生产使用;但 README 对模型调用方式、任务与输出产物描述具体,仓库公开可核验。

原始记录:zjcolacode 用 Qwen3.6 Plus 视觉模型构建视频精华自动剪辑智能体

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zjkaaap / GradeLens 使用 Qwen3.6 Plus 处理文档理解和结构化处理

zjkaaap / GradeLens · Qwen3.6 Plus

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Plus 来源平台:github 最后复核:2026-06-27T08:58:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zjkaaap / GradeLens 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T08:58:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:教师上传 .docx 试卷后系统解析题干、标准答案与分值;学生上传作答照片后,多模态模型直接读取手写作答并对照答案评分。

公开产物:公开 README 说明 GradeLens 可输出结构化得分、满分、扣分点和评语,支持整卷评分;项目徽章和简介均明确标注模型为 qwen3.6-plus。

模型作用:Qwen3.6 Plus 直接执行看图、理解手写内容、对照标准答案评分和生成评语的核心流程,避免传统 OCR 加文本模型两段式方案对手写公式的识别损失。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为项目 README 自述,未见外部学校/机构部署证明;仓库公开,模型、任务、输出和组织信息明确。

原始记录:GradeLens 智阅用 Qwen3.6 Plus 多模态能力实现端到端自动阅卷

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

IvanLark 使用 GPT-5.4 (xhigh) 处理软件工程任务执行

IvanLark · GPT-5.4 (xhigh)

A
厂商:OpenAI 模型:GPT-5.4 (xhigh) 来源平台:github 最后复核:2026-06-27T08:59:16Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

IvanLark 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T08:59:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Non-invasive OpenClaw runtime shim that enables gpt-5.4 plus /think xhigh, hooks OpenClaw and bundled provider xhigh capability checks, and optionally injects priority service tier into the OpenAI Responses payload path.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository provides a shim, installation/runtime instructions, compatibility notes for OpenClaw 2026.3.2, verification commands, rollback steps, and Chinese/English documentation.

模型作用:GPT-5.4 (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.4 xhigh is the target model/effort combination being enabled; the shim preserves xhigh through OpenClaw capability checks and downstream provider handling so the configured agent can actually use that reasoning mo…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is an integration/enablement artifact rather than an end-user application; README requires the user's backend/provider to already support gpt-5.4 with xhigh.

原始记录:IvanLark built an OpenClaw shim to keep gpt-5.4 xhigh enabled through provider capability checks

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chenzinuo2005 / music_creator 使用 Qwen3.6 Plus 处理多模态内容处理

chenzinuo2005 / music_creator · Qwen3.6 Plus

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Plus 来源平台:github 最后复核:2026-06-27T08:58:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chenzinuo2005 / music_creator 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T08:58:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:构建视频理解加音乐生成工作流:加载视频文件并以 Base64 传入 Qwen3.6-plus,提取 mood、energy、style、instruments、description 等音乐参数,再转成 MiniMax Music2.6 的生成提示词。

公开产物:公开仓库 README 展示 GUI/FastAPI 双入口、SSE 流式视频分析、音频预览/下载等功能;工作流程明确说明 Qwen3.6-plus 输出结构化 JSON 后用于生成 MP3 背景音乐。

模型作用:Qwen3.6 Plus 承担视频内容理解与音乐参数抽取,把视频画面/内容转化为可供音乐模型消费的结构化创作意图。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:案例来自开源项目说明,产物为工具代码而非已验证商业产品;但模型绑定、任务链路和输出结果均可在 README 核验。

原始记录:music_creator 用 Qwen3.6 Plus 分析视频并驱动 MiniMax Music2.6 生成背景音乐

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MM1ng / ProcuraAI 使用 Qwen3.7 Max 处理知识检索和问答

MM1ng / ProcuraAI · Qwen3.7 Max

A
厂商:Qwen / Alibaba 模型:Qwen3.7 Max 来源平台:github 最后复核:2026-06-27T09:00:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MM1ng / ProcuraAI 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T09:00:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ProcuraAI 是一个企业采购助手,用自然语言理解采购需求,检索商品,生成采购计划,并创建订单与支付流程;README 明确写出支持 Qwen3.7-Max 与 Qwen-Max,并在环境变量中使用 QWEN_MODEL=qwen3.7-max。

公开产物:公开仓库提供了完整的前后端采购 Agent、RAG 商品检索、采购计划接口、聊天界面、订单/Stripe 测试支付流程,以及示例请求:为 20 名实习生在预算内采购键盘、鼠标和耳机。

模型作用:Qwen3.7-Max 作为 LangChain Tongyi / Alibaba Cloud 接入的 LLM,用于解析自然语言采购意图、生成采购推荐与计划,并支撑多轮采购对话。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:社区 GitHub 项目;证据来自 README 和配置文件,未见独立生产客户故事,但仓库包含可运行产物和明确模型绑定。

原始记录:ProcuraAI 用 Qwen3.7-Max 构建企业采购 Agent

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

AugustOsei 使用 Qwen3.7 Max 处理智能体流程编排

AugustOsei · Qwen3.7 Max

A
厂商:Qwen / Alibaba 模型:Qwen3.7 Max 来源平台:github 最后复核:2026-06-27T09:00:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AugustOsei 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T09:00:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开发者公开说明该 WebGL cloth simulation 项目使用 Qwen Code (qwen3.7-max) 作为 AI coding assistant 开发;项目目标是实现可交互旗帜物理、动态风、GLSL 渲染和多种旗帜样式。

公开产物:产物包括可访问的 Vercel 在线 Demo 和 GitHub 源码;页面实现实时布料物理、拖拽交互、风力预设、电影感渲染、调试网格等功能,源码中还在画布纹理写入 made with qwen3.7-max。

模型作用:qwen3.7-max 通过 Qwen Code 作为编码助手,参与生成和迭代 vanilla JavaScript / WebGL / GLSL 实现,帮助完成交互式创意开发产物。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人开源 Demo;证据明确绑定 qwen3.7-max 且在线产物可访问,但没有详细对话日志说明模型参与比例。

原始记录:AugustOsei 用 Qwen Code qwen3.7-max 开发 WebGL 交互旗帜仿真 Demo

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ada682 / hermes-openclaw 使用 Qwen3.7 Max 处理多模态内容处理

ada682 / hermes-openclaw · Qwen3.7 Max

A
厂商:Qwen / Alibaba 模型:Qwen3.7 Max 来源平台:github 最后复核:2026-06-27T09:00:28Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ada682 / hermes-openclaw 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T09:00:28Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:hermes-openclaw 是一个面向 Hermes / OpenClaw Agent 的本地代理项目;代码中将 QWEN_MODEL 默认值设为 qwen3.7-max,并把 Qwen 对话流转换为 OpenAI-compatible SSE、tools、multimodal message 和配置入口。

公开产物:公开仓库提供 reverse-proxy.js、start.js 与配置修补脚本,可启动 qwen/kimi/deepseek 代理并把本地 Agent 的 OPENAI_BASE_URL 指向代理;README 说明其用于 agents using qwen / kimi / deepseek model。

模型作用:qwen3.7-max 是 Qwen 代理默认文本模型,负责本地 Agent 的聊天补全、推理内容转发和工具调用输出,项目围绕把该模型适配到 OpenAI API 兼容工作流。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:社区逆向代理项目,可能涉及非官方接入方式;作为公开开发工具案例可核验,但不应作为官方合规客户案例展示。

原始记录:hermes-openclaw 将 qwen3.7-max 接入 OpenAI-compatible 本地代理

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

TTAWDTT 使用 GPT-5.2 (xhigh) 处理软件工程任务执行

TTAWDTT · GPT-5.2 (xhigh)

A
厂商:OpenAI 模型:GPT-5.2 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:02:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

TTAWDTT 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:02:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:仓库说明明确称同济大学 2025 数据结构课程设计作业答案采用 gpt-5.2-xhigh + Codex 生成,公开仓库保存了课程设计代码/答案产物。

公开产物:公开 GitHub 仓库包含可访问的课程设计作业答案与代码文件,供他人下载/复用。

模型作用:GPT-5.2 xhigh 通过 Codex 参与生成数据结构课程设计答案和代码实现。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:用途涉及课程作业代写/抄袭风险;证据来自仓库公开描述与可访问产物,不作为合规背书。

原始记录:TTAWDTT 用 GPT-5.2 xhigh + Codex 生成同济软件工程数据结构课程设计答案仓库

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

j-kemble 使用 GPT-5.2 (xhigh) 处理软件工程任务执行

j-kemble · GPT-5.2 (xhigh)

A
厂商:OpenAI 模型:GPT-5.2 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:02:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

j-kemble 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:02:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:仓库公开描述称该本地凭证生成与存储工具由 AI 构建,明确列出 Warp Terminal/GPT 5.2 Codex xhigh reasoning、Claude 4.5 Haiku 与 Gemini-CLI;README 展示了可安装运行的 TUI、本地加密 vault、CSV 导入导出等功能。

公开产物:产物是可通过 PyPI 或源码安装运行的 Generate It CLI/TUI,用于生成随机密码、passphrase、用户名并在本地加密 vault 中管理凭证。

模型作用:GPT-5.2 Codex xhigh reasoning 作为协作编码模型之一,参与实现本地凭证生成器、TUI 与加密凭证库等功能。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目同时声明使用了 Claude 与 Gemini-CLI,单项贡献比例无法从公开证据精确拆分;凭证管理工具需用户自行安全审计。

原始记录:j-kemble 用 GPT-5.2 Codex xhigh reasoning 等模型构建本地凭证生成与加密管理工具 Generate It

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Healthy-Nebula-3603 使用 GPT-5.2 (xhigh) 处理可玩交互原型构建

Healthy-Nebula-3603 · GPT-5.2 (xhigh)

A
厂商:OpenAI 模型:GPT-5.2 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:02:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Healthy-Nebula-3603 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T09:02:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:公开仓库标题/URL 标注 gpt5.2-codex_xhigh proof of concept GBA emulator in assembly;README 描述让 Codex 构建可运行 Nintendo GBA 模拟器,用 x86-64 assembly 实现核心 emulation subsystem,并以 SuperMarioAdvance 作为验收目标。

公开产物:产物是公开 GitHub 仓库中的 GBA emulator proof-of-concept,包括汇编核心、SDL2 前端/宿主层设计和构建运行说明。

模型作用:GPT-5.2 Codex xhigh 被用于规划、编码、测试截图/自玩和调试模拟器,协助完成从提示到可运行原型的系统编程任务。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:仓库 README 正文另有 GPT 5.3 Codex xhigh 字样,与仓库标题/URL 的 GPT 5.2 Codex xhigh 存在不一致;此条按公开仓库标题/URL 绑定,需后续复核作者实际使用的模型版本。

原始记录:Healthy-Nebula-3603 用 GPT-5.2 Codex xhigh 生成纯汇编 GBA 模拟器原型

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

KaishuShito 使用 GPT-5.2 (xhigh) 处理浏览器 3D 世界构建

KaishuShito · GPT-5.2 (xhigh)

A
厂商:OpenAI 模型:GPT-5.2 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:02:39Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

KaishuShito 公开的3D 与 Web 交互案例,来源为 公开代码库,复核于 2026-06-27T09:02:39Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:仓库公开描述列出 GPT-5.2 xhigh;README 给出相同提示下的 OpenAI 输出文件,包括 LP Design、SVG Animation、3D Game、Creative Writing 和 Philosophy 等任务。

公开产物:公开仓库保存了 outputs/openai 下的 HTML、SVG/动画/3D 游戏和 Markdown 写作产物,可直接查看模型输出。

模型作用:GPT-5.2 xhigh 用于根据固定提示生成前端页面、动画、游戏和文本创作结果,形成可对照的公开产物。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该项目是模型输出对比而非生产部署;README 的模型列表也提到 GPT-5.3 Codex,需以后续仓库描述/输出元数据复核具体 OpenAI 输出对应版本。

原始记录:KaishuShito 在 VibeCheck 中用 GPT-5.2 xhigh 生成网页、SVG、3D 游戏与写作输出用于模型对比

已有真实案例 3D 与 Web 交互公开代码库A 类可核验real_case auto_approved 进入模型卡精选

dakshjain-1616 使用 Grok 4.20 0309 处理研究分析和报告生成

dakshjain-1616 · Grok 4.20 0309

A
厂商:xAI / Grok 模型:Grok 4.20 0309 来源平台:GitHub 最后复核:2026-06-27T09:08:24Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

dakshjain-1616 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T09:08:24Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:dakshjain-1616 发布 Python 多智能体模拟环境,README 明确说明项目通过 OpenRouter 使用 x-ai/grok-4.20-multi-agent-beta,按 Researcher、Analyst、Synthesizer 三个专门代理顺序处理用户任务,并提供 rich/plotext 实时 CLI 仪表盘。

公开产物:公开仓库提供可运行代码、配置和 CLI 流程:Researcher 收集事实,Analyst 提炼模式/风险/机会,Synthesizer 生成最终整合答案,同时保存 run_log.json、Rich 终端输出和最终仪表盘图表。

模型作用:Grok 4.20 Multi-Agent 作为三个代理调用的核心推理模型,驱动事实收集、分析和综合输出,是该模拟器多阶段任务处理能力的主要模型来源。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据使用 OpenRouter beta 别名 x-ai/grok-4.20-multi-agent-beta;名称未带 0309 后缀,但可绑定到 Grok 4.20 multi-agent/0309 家族。

原始记录:grok-multiagent-simulator 用 Grok 4.20 Multi-Agent 串联 Researcher/Analyst/Synthesizer

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

PrashanthBhaskara / ProphetHacks 使用 Grok 4.20 0309 处理真实任务执行

PrashanthBhaskara / ProphetHacks · Grok 4.20 0309

A
厂商:xAI / Grok 模型:Grok 4.20 0309 来源平台:GitHub 最后复核:2026-06-27T09:08:24Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

PrashanthBhaskara / ProphetHacks 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T09:08:24Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ProphetHacks 在预测市场系统中实现 Grok-via-OpenRouter forecaster,代码注释明确称其调用 Grok-4.20 through OpenRouter,并为 Kalshi-style markets 设计 trust-extreme calibration prompt、二向提示和噪声过滤。配置 FINAL.json 中 grok_lane 的模型为 x-ai/grok-4.20。

公开产物:公开代码记录该 Grok forecaster 在 2026 Sports holdout 上通过噪声过滤组合取得 Brier 0.203,对比市场基线 0.212,并作为集成预测管线的一路模型输出概率预测。

模型作用:Grok 4.20 负责生成预测市场的概率判断;系统围绕其校准弱点加入 trust-extreme 提示、P(YES)/P(NO) 双向询问和过滤策略,使其预测可被集成器加权使用。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是公开仓库中的应用代码与验证说明,不是官方客户故事;模型以 OpenRouter x-ai/grok-4.20 形式出现,可绑定 Grok 4.20 0309 家族。

原始记录:ProphetHacks 将 Grok 4.20 作为 Kalshi/预测市场集成预测器的一路模型

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kvsdileep 使用 Grok 4.20 0309 处理真实任务执行

kvsdileep · Grok 4.20 0309

A
厂商:xAI / Grok 模型:Grok 4.20 0309 来源平台:GitHub 最后复核:2026-06-27T09:08:24Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kvsdileep 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T09:08:24Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:kvsdileep 的 Model Council 是本地 Web App,让四个模型并行回答、互相批评,再由综合模型调和。README 描述 Grok 作为 council member;配置文件 config/models.ts 将 grok 槽位设为 x-ai/grok-4.20,并说明该 ID 经 OpenRouter live catalog probe 返回 HTTP 200 后替换原设计 ID。

公开产物:项目实现了 POST SSE 流式接口、四列模型输出、同伴批评和最终综合的前端/后端结构;仓库计划与验证文档记录了 curl smoke 产生 4 个 distinct delta model streams,其中包括 Grok 槽位。

模型作用:Grok 4.20 作为四个委员会成员之一提供独立答案和批评视角,参与多模型辩论阶段,为最终综合答案提供 xAI/Grok 侧的推理输入。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 中模型表曾写 Grok 4.2 设计稿,但实际 config/models.ts 与验证文档明确将 live ID 替换为 x-ai/grok-4.20;作为真实开发中模型接入案例收录。

原始记录:Model Council 用 Grok 4.20 参与多模型辩论和同伴评审

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

owjk123 使用 Grok 4.20 0309 处理真实任务执行

owjk123 · Grok 4.20 0309

A
厂商:xAI / Grok 模型:Grok 4.20 0309 来源平台:GitHub 最后复核:2026-06-27T10:20:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

owjk123 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T10:20:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:GrokNovelApp 是一个基于 Grok 4.20 API 的 Android 小说生成应用,README 明确描述其使用 Grok 4.20 API 生成精彩小说内容,支持奇幻、科幻、爱情、悬疑、仙侠、武侠等类型,以及互动剧情选择和本地历史保存。

公开产物:公开仓库提供 Kotlin / Jetpack Compose Android 应用源码、项目结构、构建说明和 Material3 界面设计,目标产物是可构建的 AI 小说生成 Android 应用。

模型作用:Grok 4.20 API 用于根据用户选择和剧情上下文生成小说正文与后续分支,是该应用创意文本生成和互动剧情推进的核心模型能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据明确写 Grok 4.20 API,但未带 0309 后缀;按 Grok 4.20 家族映射到 Grok 4.20 0309 任务模型。仓库和 README 均可公开访问,属于具体应用产物而非教程、benchmark 或综述。

原始记录:owjk123 的 GrokNovelApp 使用 Grok 4.20 API 生成互动小说 Android 应用

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

coasty-ai / Coasty AI 使用 Claude Mythos 5 处理智能体流程编排

coasty-ai / Coasty AI · Claude Mythos 5

A
厂商:Anthropic / Claude 模型:Claude Mythos 5 来源平台:github 最后复核:2026-06-27T10:32:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

coasty-ai / Coasty AI 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T10:32:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Coasty AI 发布 open-cowork,一个开源、跨平台的 agentic coworker,可观察屏幕并执行桌面、云 VM 或浏览器任务,支持实时步骤流、人工审批和成本控制;GitHub 仓库描述明确称该项目 built by Claude Mythos 5。

公开产物:公开产物为 TypeScript 开源仓库 coasty-ai/open-cowork,README 提供本地运行、demo GIF、BYOK/Coasty API 配置、桌面自动化和浏览器自动化说明;GitHub 搜索结果显示仓库公开、创建于 2026-06-11、包含产品页 https://coasty.ai/?view=developers。

模型作用:Claude Mythos 5 被明确标注为构建该开源 AI coworker 项目的模型,贡献于将屏幕理解、任务委派、浏览器/桌面操作、实时转录与审批流程整合成可运行的开发者产物。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 仓库描述和 README;未发现独立第三方复核其内部开发过程,因此按公开项目自述收录,不使用 benchmark、教程或集合页冒充案例。

原始记录:Coasty AI open-cowork:用 Claude Mythos 5 构建开源 AI coworker 桌面/浏览器自动化项目

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Thibaultbm / Sorank 使用 Claude Mythos 5 处理软件工程任务执行

Thibaultbm / Sorank · Claude Mythos 5

A
厂商:Anthropic / Claude 模型:Claude Mythos 5 来源平台:github 最后复核:2026-06-27T10:32:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Thibaultbm / Sorank 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:32:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Thibaultbm 发布 claude-seo-geo,一个面向 Claude Code 的 SEO 与 GEO 技能包,README 明确写明“SEO & GEO skills for Claude Code, built with Claude Mythos 5”,覆盖技术审计、外链策略、AI 优化内容、本地可见性和社交放大等任务。

公开产物:公开产物为 GitHub 仓库 Thibaultbm/claude-seo-geo,README 展示 15 个零依赖技能、Claude Code plugin 安装命令、skills CLI 安装命令,并说明这些技能基于 115+ 真实 agency audit calls 与 2026 年 AI search 证据整理。

模型作用:Claude Mythos 5 被明确标注为构建该 Claude Code SEO/GEO 技能包的模型,贡献于把 SEO/GEO 工作流拆解为可安装技能,并形成可在 Obsidian/Claude Code 中执行的审计、内容与增长任务模板。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README 和仓库元数据;这是开源产物自述的 built-with 案例,未把泛化 SEO 教程、benchmark 或新闻综述计入。

原始记录:Claude SEO GEO:用 Claude Mythos 5 构建面向 Claude Code 的 SEO/GEO 技能包

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

frahlg 使用 GPT-5.5 (xhigh) 处理软件工程任务执行

frahlg · GPT-5.5 (xhigh)

A
厂商:OpenAI 模型:GPT-5.5 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:14:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

frahlg 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Fusion 是一个 Claude Code skill,把高风险问题并行发送给多个前沿模型,由 GPT-5.5 (xhigh) 作为其中一个 blind panelist 进行独立推理,再由 Opus 4.8 汇总裁决。

公开产物:公开仓库提供了可安装的 Fusion skill、README 流程图和 assets/fusion-demo.gif 演示,用于输出带 verdict、consensus、contradictions、blind spots 与 audit trail 的最终答案。

模型作用:GPT-5.5 (xhigh) 在 panel 阶段提供独立推理路径和候选答案,贡献多模型分歧与互证信号,帮助最终答案避免单模型盲区。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README;它说明模型在工具中的设计用途和演示产物,但不提供生产使用指标。

原始记录:Fusion 用 GPT-5.5 xhigh 作为盲审 panelist 生成可审计答案

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kirillshsh 使用 GPT-5.5 (xhigh) 处理软件工程任务执行

kirillshsh · GPT-5.5 (xhigh)

A
厂商:OpenAI 模型:GPT-5.5 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:14:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kirillshsh 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:shimplify 是全局 Codex skill,用三个专门的 GPT-5.5 xhigh Codex subagents 检查已能运行但存在 AI-generated clutter 的代码,包括重复 helper、膨胀参数、无必要抽象、重复检查和低效热路径。

公开产物:公开仓库提供 npm/脚本安装方式、reuse/quality/efficiency 三类 agent 配置,以及在当前分支应用 compact fixes 的工作流说明。

模型作用:GPT-5.5 (xhigh) 分别承担 reuse、quality、efficiency 三个只读 reviewer 角色,提出可执行的复用、简化和性能改进发现,再由主 Codex agent 聚合并应用。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 公开说明模型角色和产物;没有公开量化的缺陷修复成功率。

原始记录:shimplify 用三个 GPT-5.5 xhigh Codex 子代理清理 AI 代码冗余

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

phiraml 使用 GPT-5.5 (xhigh) 处理软件工程任务执行

phiraml · GPT-5.5 (xhigh)

A
厂商:OpenAI 模型:GPT-5.5 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:14:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

phiraml 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:claude-codex-judge 是 Claude Code skill,README 明确说明它使用 OpenAI codex CLI 的 gpt-5.5、xhigh reasoning,作为独立只读 judging agent 审查 Claude 的 plans、designs、diffs 和 decisions。

公开产物:公开仓库包含 .claude/skills/codex-judge/SKILL.md、codex-judge.sh wrapper 和使用示例;输出 verdict、关键修复建议,并支持跨轮 session continuity。

模型作用:GPT-5.5 (xhigh) 负责读取代码库上下文并提出反驳、风险和判断结论,为 Claude 的计划或 diff 提供第二模型审查。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自项目 README 和公开脚本;案例是开源工具而非客户生产故事。

原始记录:claude-codex-judge 用 GPT-5.5 xhigh 作为只读判断代理

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Sol3Coder 使用 GPT-5.5 (xhigh) 处理智能体流程编排

Sol3Coder · GPT-5.5 (xhigh)

A
厂商:OpenAI 模型:GPT-5.5 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:14:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Sol3Coder 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T09:14:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:该仓库公开了面向 MISRA C++:2023 的通用 AI agent skill,用于 C++ review、新代码指导和启发式 safety gate checks;README 明确写明该 skill was created by GPT-5.5 with xhigh reasoning effort。

公开产物:公开产物包含 .agents/skills/misra-cpp-2023/SKILL.md、review workflow、coding guidance、rule index、scan_cpp_misra.py 启发式扫描器和 unittest 验证入口。

模型作用:GPT-5.5 (xhigh) 负责生成该 MISRA C++:2023 agent skill 的结构、审查流程、指导文档与辅助扫描脚本。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:项目注明 scanner 不是认证 MISRA checker,且因版权限制不包含原始 MISRA PDF;证据充分绑定模型创建产物,但不是生产客户案例。

原始记录:MISRA C++:2023 Agent Skill 由 GPT-5.5 xhigh 创建用于 C++ 安全审查

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Z.ai / 智谱AI 使用 GLM-5 处理软件工程任务执行

Z.ai / 智谱AI · GLM-5

A
厂商:Z AI / GLM 模型:GLM-5 来源平台:official_web 最后复核:2026-06-27T09:14:54Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Z.ai / 智谱AI 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T09:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:GLM Coding Plan packages GLM-5 family models into an AI coding subscription for natural-language programming, code debugging and fixing, codebase Q&A, automated lint/merge-conflict/release-note tasks, and integrations w…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public subscription/product page and official documentation are live; the docs list supported coding scenarios, tool integrations, usage/quota management, and available GLM-5 family models, with historical GLM-5/GLM-5…

模型作用:GLM-5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5 family models provide the reasoning, coding, tool-use, and agentic execution layer behind the coding-plan workflows, turning natural-language developer requests into plans, code changes, debugging suggestions, and…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Official product evidence; current docs state historical GLM-5/GLM-5.1 calls are automatically switched to GLM-5.2, so this is a GLM-5 family case rather than proof of continued serving of the original GLM-5 endpoint.

原始记录:Z.ai GLM Coding Plan uses GLM-5 family models for AI coding workflows

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

ltmoerdani 使用 GLM-5 处理软件工程任务执行

ltmoerdani · GLM-5

A
厂商:Z AI / GLM 模型:GLM-5 来源平台:github 最后复核:2026-06-27T09:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ltmoerdani 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A VS Code extension registers Z.AI GLM series models, explicitly including GLM-5.2, GLM-5.1, GLM-5, and GLM-5-Turbo, into GitHub Copilot Chat through the VS Code Language Model Chat Provider API so developers can select…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains installation/configuration documentation, a model support table for GLM-5 family variants, and an extension implementation intended to make Copilot Chat operate with Z.AI GLM models instea…

模型作用:GLM-5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5 supplies the chat, coding, and agentic planning responses surfaced through Copilot Chat; the extension contributes the VS Code integration and credential/model-provider plumbing.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Open-source extension evidence; actual user deployment depends on users providing a valid Z.AI key and the extension being installed in VS Code.

原始记录:Z.AI for GitHub Copilot Chat exposes GLM-5 models inside VS Code Copilot Chat

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

delta-whiplash 使用 GLM-5 处理软件工程任务执行

delta-whiplash · GLM-5

A
厂商:Z AI / GLM 模型:GLM-5 来源平台:github 最后复核:2026-06-27T09:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

delta-whiplash 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:An unofficial desktop application packages the Z.ai chatbot/agent web experience, described in the README as powered by GLM-5 and GLM-4.7, into lightweight Linux, macOS, and Windows desktop builds with system-tray, nati…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo provides source code, build instructions, and release-oriented packaging for a desktop client that opens chat.z.ai as a native-feeling app across operating systems.

模型作用:GLM-5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5 is the model family powering the underlying Z.ai chatbot/agent interactions; the desktop wrapper makes those model capabilities available in a native desktop workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Unofficial wrapper around Z.ai; evidence binds the app to the GLM-5-powered Z.ai service but does not expose private runtime logs from end users.

原始记录:Z.ai Desktop wraps the GLM-5-powered Z.ai chat product as a cross-platform desktop app

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xiaoY233 使用 GLM-5 处理软件工程任务执行

xiaoY233 · GLM-5

A
厂商:Z AI / GLM 模型:GLM-5 来源平台:github 最后复核:2026-06-27T09:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xiaoY233 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Chat2API is a multi-provider AI service management tool that connects official web UIs, including GLM and Z.ai, and exposes them through OpenAI-compatible API endpoints for clients such as OpenClaw, Cline, Roo-Code, Che…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents provider support for GLM/Z.ai, dashboard monitoring, API key management, request logging, model mapping, context management, and function-calling support; its supported-provider table includes G…

模型作用:GLM-5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5 family models provide the upstream chat/coding/model responses, while Chat2API contributes account/session handling, model mapping, monitoring, and OpenAI-compatible serving for downstream tools.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The README names GLM-5.1 rather than the original GLM-5 row; accepted as a GLM-5-family integration. It is an open-source integration tool, not a vendor customer story.

原始记录:Chat2API includes GLM/Z.ai support to expose GLM-5 family chat through OpenAI-compatible APIs

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

MisterKarott 使用 GLM-5 处理软件工程任务执行

MisterKarott · GLM-5

A
厂商:Z AI / GLM 模型:GLM-5 来源平台:github 最后复核:2026-06-27T09:14:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

MisterKarott 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:14:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:glm-quota is a Claude Code plugin that activates when developers run Claude Code with Z.ai/GLM, displaying the current model, context-window usage, token counts, compacting reminders, Z.ai quota windows, and MCP tool-ca…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo documents installation through a Claude plugin marketplace or manual clone and shows a statusline example for a GLM-5.1 run with 1M context, context percentage, token counts, quota reset timers, and MCP …

模型作用:GLM-5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5 family models are the coding backend whose context and quota usage the plugin observes; the plugin turns model-runtime limits into actionable developer feedback during coding sessions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence specifically shows GLM-5.1 in the statusline example; treated as GLM-5-family evidence. The plugin observes usage rather than generating end-user content itself.

原始记录:glm-quota tracks GLM-5 family usage and context inside Claude Code statuslines

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

owjk123 / NovelForge 使用 Grok 4.3 (high) 处理软件工程任务执行

owjk123 / NovelForge · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:17:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

owjk123 / NovelForge 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:NovelForge is an Android novel-writing application that connects to the Grok 4.3 API to help authors generate medium- and long-form fiction, including one-click chapter generation, multi-genre support, chapter managemen…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains the Android app implementation and README describing generated novel chapters, saved chapter history, continuation support, local Room storage, and a buildable APK workflow.

模型作用:Grok 4.3 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 4.3 is the text generation engine used to produce and continue chapter content from author inputs while maintaining story context across chapters.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence names Grok 4.3 API rather than spelling out the high reasoning tier; accepted as the public Grok 4.3 family/API binding for this task model.

原始记录:NovelForge uses Grok 4.3 API to generate and continue long-form novel chapters in an Android app

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

kivo360 / OmoiOS 使用 GLM-5.1 处理软件工程任务执行

kivo360 / OmoiOS · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:GitHub 最后复核:2026-06-27T09:17:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kivo360 / OmoiOS 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OmoiOS is a self-hostable system where users describe software work and the platform plans a dependency graph, runs a self-supervising swarm of coding agents in isolated sandboxes, and lands pull requests. Its sandboxed…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides an executable OmoiOS artifact for sandboxed coding-agent workflows; its README states the system can clone projects, run test suites in spawned sandboxes, and drive PR-oriented coding sess…

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 is bound as the default LLM backend for the sandboxed coding agent, supplying the reasoning and code-generation capability used inside OmoiOS agent sessions unless the operator selects another model.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from repository code plus README; no independent production deployment metrics are provided.

原始记录:OmoiOS uses GLM-5.1 as the default model for sandboxed coding-agent execution

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

colesmcintosh / Meme Volatility Ind… 使用 Grok 4.3 (high) 处理软件工程任务执行

colesmcintosh / Meme Volatility Index · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:17:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

colesmcintosh / Meme Volatility Index 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Meme Volatility Index is a real-time meme-stock ticker for X where a user enters a topic and the app analyzes live social posts to compute hype score, momentum, chart updates, AI captions, key insights, and top signal p…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository README documents a Next.js/Vercel AI SDK app that outputs a 1-10 hype score, Pumping/Dumping/Stable momentum, a live chart, AI meme caption, key insights, and per-post hype ratings.

模型作用:Grok 4.3 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The app calls OpenRouter model slug x-ai/grok-4.3 with an xSearch/web plugin and structured generateObject output, so Grok 4.3 performs the live X-post interpretation and structured hype scoring.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence uses the OpenRouter slug x-ai/grok-4.3; it does not separately label the tier as high.

原始记录:Meme Volatility Index uses x-ai/grok-4.3 plus xSearch to score live meme-stock hype from X posts

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

wxy2ab / akinterpreter 使用 GLM-5.1 处理软件工程任务执行

wxy2ab / akinterpreter · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:GitHub 最后复核:2026-06-27T09:17:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

wxy2ab / akinterpreter 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:akinterpreter is an LLM-powered financial market query and analysis tool that retrieves market data, generates analysis code, executes plans, corrects errors, and produces reports from natural-language requests. Its GLM…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repo ships a CLI/web financial-analysis artifact that can query market data via providers such as akshare and tushare and return usable analysis results and reports from natural-language prompts.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 provides the language-model layer for the GLM client path, contributing planning, code generation, error correction, and report generation for financial data analysis workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence confirms the default GLM client model and the repository artifact; deployment/customer usage volume is not disclosed.

原始记录:akinterpreter defaults its GLM client to GLM-5.1 for financial market analysis

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

jmtroller 使用 Grok 4.3 (high) 处理文档理解和结构化处理

jmtroller · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:19:45Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

jmtroller 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T09:19:45Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:从破产案件 PACER PDF 文档中抽取结构化字段;README 明确写明默认 provider 是 Grok 4.3(xAI Responses API),并提供 doc-extraction extract 命令对 docketId 对应 PDF 自动识别文档类型后抽取。

公开产物:产物是一个可运行的 Python 服务,输出写入 MongoDB BankruptcyIntel 集合;每个抽取字段采用 {value, source_page} 格式,并带 provenance、raw、normalized JSON,可进一步投影给 Laravel/MCP/REST 使用。

模型作用:Grok 4.3 作为默认抽取模型处理 PDF 内容理解、文档类型相关字段识别、结构化 JSON 生成和来源页码定位;Anthropic/OpenAI 仅作为可替换备用 provider。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:公开证据来自仓库 README;未看到线上生产数据样例,案例粒度按公开开源服务/工程产物采集。

原始记录:jmtroller 用 Grok 4.3 为破产案件 PACER PDF 做结构化抽取

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CheshireMew / MediaFlow 使用 GLM-5.1 处理多模态内容处理

CheshireMew / MediaFlow · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:GitHub 最后复核:2026-06-27T09:17:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CheshireMew / MediaFlow 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T09:17:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:MediaFlow is a local video subtitle generation and processing workstation covering video download, Whisper-based transcription, subtitle translation, glossary-aware editing, preview, and video composition. Its frontend …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public Electron + React + FastAPI application provides an end-to-end artifact for generating, translating, editing, and composing subtitles for local videos.

模型作用:GLM-5.1 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 is the default configured Zhipu LLM for translation-related workflow steps, supporting subtitle translation and language-processing tasks in the MediaFlow workstation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public configuration binding plus project README; individual translated video outputs are not linked.

原始记录:MediaFlow configures GLM-5.1 as a default Zhipu model for local video subtitle workflows

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

smarvela / Lutheran Confessional Q&A 使用 Grok 4.3 (high) 处理代码审查和测试生成

smarvela / Lutheran Confessional Q&A · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:17:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

smarvela / Lutheran Confessional Q&A 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T09:17:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Lutheran Confessional Q&A is a chat-based web application that answers theological questions using retrieval-augmented generation over the Book of Concord / Tunnustuskirjat, Finnish confessional texts, and public-domain…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents a modern chat interface whose answers are explicitly grounded in retrieved passages, show sources for transparency, and work in English or Finnish depending on ingested material.

模型作用:Grok 4.3 (high) 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Grok 4.3 generates the grounded answers after retrieval over a local vector database; the README says the model only sees retrieved passages.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence names Grok 4.3, not an explicit high tier; accepted as family binding to the task model.

原始记录:Lutheran Confessional Q&A uses Grok 4.3 with RAG to answer questions grounded in confessional texts and the KJV Bible

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

wangyingqi-x / workflow-showcase-pl… 使用 GLM-5.1 处理软件工程任务执行

wangyingqi-x / workflow-showcase-platform · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:GitHub 最后复核:2026-06-27T09:17:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

wangyingqi-x / workflow-showcase-platform 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:workflow-showcase-platform is a public Java workflow and agent runtime demo with DAG orchestration, bounded ReAct loops, dynamic MCP import, human-in-the-loop resume, trace/timeline/variables display, and an Agent Studi…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository publishes a runnable showcase platform where the frontend drives real backend workflow execution, tool routing, ReAct loops, and resumable human confirmation rather than a static mock UI.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 supplies the default/fallback chat model for agent nodes and ReAct execution in the workflow engine, enabling tool-using reasoning and natural-language control of workflows.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from source configuration and README; no external user deployment is claimed.

原始记录:workflow-showcase-platform uses GLM-5.1 as the fallback model for a Java agent/workflow runtime

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Phazzie / SkepticalWombat 使用 Grok 4.3 (high) 处理软件工程任务执行

Phazzie / SkepticalWombat · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:17:13Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Phazzie / SkepticalWombat 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:13Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SkepticalWombat is an AI writing-coach web app where users dump drafts and the app reads them, detects skipped gaps, calls out contradictions, helps structure chapters or beats, and supports persistent project chat.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository documents Brain Dump, Wombat Reads, Structure It, and Talk It Through features, plus a Next.js/Neon/Stack Auth implementation for the writing-coach product.

模型作用:Grok 4.3 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Grok 4.3 powers the text-analysis and coaching layer, specifically gap detection, contradiction spotting, auto-structuring, and context-aware writing feedback.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence identifies xAI Grok 4.3 as the AI text layer but does not explicitly state high tier; accepted as task-family binding.

原始记录:SkepticalWombat uses Grok 4.3 as a writing coach for gap detection, contradiction spotting, and story structuring

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenSquilla 使用 GLM-5.1 处理软件工程任务执行

OpenSquilla · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:GitHub 最后复核:2026-06-27T09:17:37Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenSquilla 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:17:37Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:OpenSquilla is a token-efficient microkernel AI agent for CLI, Web UI, and chat channels with a local model router, persistent memory, layered sandbox, web search, and embeddings. Its gateway configuration maps Zhipu 's…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public OpenSquilla repository ships an agent artifact with CLI/Web/chat entry points and routing profiles that select models by task tier and budget.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 is used as the Zhipu model for complex text and high-reasoning routes, contributing stronger reasoning/text handling within OpenSquilla's router-managed agent loop.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence confirms routing configuration in the public product repository; production traffic and outcome metrics are not published.

原始记录:OpenSquilla routes Zhipu complex-text tiers to GLM-5.1

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Prateek Rungta / prateek 使用 GPT-5.2 (xhigh) 处理软件工程任务执行

Prateek Rungta / prateek · GPT-5.2 (xhigh)

A
厂商:OpenAI 模型:GPT-5.2 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:22:25Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Prateek Rungta / prateek 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:25Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:公开 PR 描述称这是一次实验:观察 Ralph 搭配 gpt-5.2-xhigh 处理 Anthropic performance take-home 的效果;PR 分支名为 GPT-5.2-xhigh,评论记录了多轮 Ralph iteration、Codex exit、提交和测试结果。

公开产物:公开 PR/分支包含 56 个提交和可访问代码改动;迭代记录显示模型驱动的优化从 baseline 147734 cycles 降到 98583 cycles,随后通过 SIMD/vectorization 降到 12369 cycles,并保留了 PRD、AGENTS.md、进度和实现文件。

模型作用:GPT-5.2 xhigh 通过 Codex/Ralph 循环参与生成 PRD/AGENTS.md、实现 greedy VLIW packer、SIMD/vectorized kernel、运行 submission tests、更新进度并提交代码,是该公开优化实验的核心编码与迭代执行模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是个人公开实验 PR,未合并且部分测量日志为本地文件;证据足以绑定模型、使用者、任务和公开代码产物,但不代表生产部署或官方背书。

原始记录:Prateek Rungta 用 Ralph + GPT-5.2-xhigh 迭代优化 Anthropic performance take-home 解法

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

thebuggybug 使用 DeepSeek-Coder-V2 处理软件工程任务执行

thebuggybug · DeepSeek-Coder-V2

A
厂商:DeepSeek 模型:DeepSeek-Coder-V2 来源平台:github 最后复核:2026-06-27T09:22:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

thebuggybug 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:thebuggybug built a VS Code extension named Deepseek Chat that integrates with the Ollama API and is currently set to use the `deepseek-coder-v2` model. The extension exposes a ChatGPT-like panel inside VS Code so devel…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains a runnable VS Code extension with README screenshots, local build/install instructions, and source code that calls `ollama.chat({ model: 'deepseek-coder-v2', ... })`. Users launch it with …

模型作用:DeepSeek-Coder-V2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek-Coder-V2 is the configured Ollama backend for the extension's conversational developer assistant. It receives user prompts from the VS Code webview and streams AI responses back into the editor chat interface, …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Small individual GitHub project rather than a large production customer story. Evidence is still concrete and public: the README names `deepseek-coder-v2`, links users to the Ollama DeepSeek-Coder-V2 model, and the repo…

原始记录:Deepseek VS Code Extension: in-editor developer chat powered by Ollama deepseek-coder-v2

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Anthropic scientists 使用 Claude Mythos 5 处理软件工程任务执行

Anthropic scientists · Claude Mythos 5

A
厂商:Anthropic / Claude 模型:Claude Mythos 5 来源平台:official_web 最后复核:2026-06-27T09:58:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Anthropic scientists 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T09:58:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Anthropic scientists used Claude Mythos 5 to produce novel molecular-biology hypotheses and compare them against Opus-class model outputs for scientific usefulness and follow-up potential.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic reported that its scientists preferred Mythos 5's molecular-biology hypotheses about 80% of the time in blinded comparisons, advanced several hypotheses to experimental evaluation, and later saw one Mythos-gen…

模型作用:Claude Mythos 5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Mythos 5 supplied the hypothesis-generation component, producing scientific mechanisms that Anthropic scientists selected for experimental evaluation and comparison against other Claude model classes.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Exact-model, first-party Anthropic evidence; details of the hypotheses and experimental program are summarized publicly but not fully disclosed. Treat as research-assistance evidence, not as a peer-reviewed result produ…

原始记录:Anthropic scientists used Claude Mythos 5 to generate molecular-biology hypotheses for experimental follow-up

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

ghostwright / Phantom 使用 GLM-5.1 处理软件工程任务执行

ghostwright / Phantom · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:github 最后复核:2026-06-27T09:22:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ghostwright / Phantom 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Phantom is an open-source AI co-worker that uses an LLM backend to operate a computer, perform coding and automation work, and can be configured to use Z.AI's Anthropic-compatible API with GLM-5.1.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository ships a working agent product with BYOM configuration; its README lists Z.AI GLM-5.1 and GLM-4.5-Air as supported provider models and describes switching providers through a YAML block.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 serves as one of Phantom's selectable reasoning/coding backends for tool-using computer-control workflows, positioned by the project as a lower-cost coding-quality alternative to Claude Opus.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public product repository and README showing explicit GLM-5.1 support; it does not publish production usage metrics or a named end-customer deployment.

原始记录:Phantom adds Z.AI GLM-5.1 as a BYOM backend for an AI co-worker that operates its own computer

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Anthropic research team 使用 Claude Mythos 5 处理软件工程任务执行

Anthropic research team · Claude Mythos 5

A
厂商:Anthropic / Claude 模型:Claude Mythos 5 来源平台:official_web 最后复核:2026-06-27T09:23:48Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Anthropic research team 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T09:23:48Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Anthropic used Claude Mythos 5 to conduct largely autonomous genomics research over more than a week: assembling single-cell data for millions of cells across 138 animal species, then designing and training a custom mac…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Anthropic reported that the trained model produced by the Mythos 5-led workflow outperformed a recent model published in Science while being 100x smaller, and that Anthropic intended to publish the results in the coming…

模型作用:Claude Mythos 5 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Mythos 5 performed long-horizon research execution with only high-level human input, combining data assembly, model design, training, and comparative analysis into a genomics research artifact.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Exact-model, first-party official evidence; the downstream paper/artifact was not yet public at collection time, so the public artifact is the Anthropic article rather than a separate dataset or publication.

原始记录:Anthropic used Claude Mythos 5 for week-long autonomous genomics research across 138 animal species

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

ginkida / Gokin 使用 GLM-5.1 处理软件工程任务执行

ginkida / Gokin · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:github 最后复核:2026-06-27T09:22:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ginkida / Gokin 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Gokin is a local CLI coding assistant that reads and edits codebases, manages context, and calls cloud or local LLM providers directly; its model options include a GLM Coding Plan tier using GLM-5.1 with a 200K context …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides installable source code and documents GLM-5.1 as a budget-friendly daily coding option, alongside provider routing, context management, and direct API-call privacy design.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 is the named model for Gokin's GLM Coding Plan, contributing long-context, low-cost code generation and agentic coding support inside the CLI workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repo README and exact GLM-5.1 mention; no independent user success metric is included.

原始记录:Gokin offers a GLM Coding Plan mode using GLM-5.1 for daily coding assistance

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

FineComputer14451 使用 Grok 4.3 (high) 处理多模态内容处理

FineComputer14451 · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:19:45Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

FineComputer14451 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T09:19:45Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Grok Imagine Cinematic Studio README 明确定位为面向 Grok Build + Grok 4.3 + Grok Imagine Video 的 23-agent 影视制作系统,可把故事转成分镜、角色 DNA、长序列、配额计划和视频提示词。

公开产物:公开产物包括 Python CLI、Streamlit Web UI、Grok plugin marketplace 清单、44 个 Grok skills、主激活提示词和模型 registry;示例命令支持用 --chat-model grok-4.3 生成视频生产提示词。

模型作用:Grok 4.3 被用作聊天/推理与创意编排模型,负责故事分析、导演/制片多 agent 协同、镜头与音画提示词规划,并与 Grok Imagine 视频模型衔接生成可执行的制作流水线。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:仓库 README 对生产能力有营销式表述;仍具备可访问代码、插件、CLI/Web UI 产物和明确模型绑定。

原始记录:FineComputer14451 用 Grok 4.3 构建多智能体影视制作工作流

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

zhayujie / CowAgent 使用 GLM-5.1 处理软件工程任务执行

zhayujie / CowAgent · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:github 最后复核:2026-06-27T09:22:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

zhayujie / CowAgent 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:CowAgent is an open-source super AI assistant and agent harness with memory, tools, skills, scheduling, browser and enterprise messaging integrations; release v2.0.7 records the addition of GLM 5.1 among new supported m…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public release and repository expose a usable assistant framework where GLM 5.1 can be selected for agentic assistant tasks, memory workflows and tool-assisted automation.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 contributes as a newly integrated LLM backend for CowAgent's assistant/agent execution, enabling users to run the product's planning, memory and tool workflows with Z AI's model family.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The evidence is an open-source release note and repository artifact; it confirms product integration but does not include quantitative deployment results.

原始记录:CowAgent v2.0.7 adds GLM 5.1 support to an open-source super assistant and agent harness

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

superagent-ai 使用 Grok 4.3 (high) 处理软件工程任务执行

superagent-ai · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:19:45Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

superagent-ai 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:19:45Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:grok-cli README 将项目描述为连接 xAI Grok API 的开源终端编码 agent,支持交互式和 headless 模式、项目目录操作、测试失败总结、重构、代码审查、子 agent、远程 Telegram 控制和批处理 API。

公开产物:公开产物是可安装的 CLI/npm 包 grok-dev;README 示例展示 grok --prompt、grok -d /path/to/repo、--verify、--format json、调度任务、媒体生成等可执行用法。

模型作用:Grok 4.3 是 README 中列出的 xAI API 支持模型之一,用于驱动编码 agent 的规划、代码阅读/修改、测试验证总结、子 agent 指令和结构化 headless 输出。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是第三方开源工具而非 xAI 官方案例;证据证明项目明确支持并面向 Grok 4.3,但具体用户运行日志未在 README 中公开。

原始记录:superagent-ai 用 Grok 4.3 打造开源终端编码 Agent

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LangcliTeam / Langcli 使用 GLM-5.1 处理软件工程任务执行

LangcliTeam / Langcli · GLM-5.1

A
厂商:Z AI / GLM 模型:GLM-5.1 来源平台:github 最后复核:2026-06-27T09:22:50Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LangcliTeam / Langcli 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:22:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Langcli is a command-line coding assistant that integrates with LangRouter so users can switch among mainstream LLMs, including GLM 5.1, during an ongoing development session without interrupting context.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides the Langcli artifact and documents GLM 5.1 as one of the models available for interactive coding sessions via LangRouter integration.

模型作用:GLM-5.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GLM-5.1 acts as a selectable coding-session model inside Langcli, contributing code-generation and reasoning capacity while preserving the active agent context across model switches.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence confirms a public tool integration and exact model name; it does not provide an external customer story or measured performance result.

原始记录:Langcli integrates GLM 5.1 through LangRouter for switchable coding sessions

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

火山引擎方舟(Volcengine Ark) 使用 Seed2.1 Pro 处理智能体流程编排

火山引擎方舟(Volcengine Ark) · Seed2.1 Pro

A
厂商:ByteDance Seed 模型:Seed2.1 Pro 来源平台:official_web 最后复核:2026-06-27T10:37:03Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

火山引擎方舟(Volcengine Ark) 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T10:37:03Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ByteDance Seed 官方发布页在体验入口中明确写明:火山方舟体验中心可选择 Doubao-Seed-2.1-Pro 或 Doubao-Seed-2.1-Turbo,且 Seed2.1 系列模型 API 已同步上线火山引擎。该案例的具体任务是把 Seed2.1 Pro 作为火山方舟模型服务/体验中心中的可选模型,供开发者通过平台试用并接入生产力、Agent、Coding 等应用工作流。

公开产物:火山方舟提供公开可访问的产品页和开发者平台入口;官方证据把 Doubao-Seed-2.1-Pro 绑定到火山方舟体验中心/API 上线,形成可核验的模型服务产物,而不是单纯的离线评测或模型发布说明。

模型作用:Doubao-Seed-2.1-Pro 在火山方舟中作为可调用的核心大模型,向开发者应用提供通用 Agent、代码工程、多模态理解、知识推理和生产力任务执行能力;平台负责模型托管、体验入口和 API 化交付。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是 ByteDance/火山引擎的一方平台集成案例,证据来自官方发布页和公开产品页;不是第三方客户故事。火山方舟控制台体验入口通常需要登录,因此使用公开产品页作为 artifact_url。

原始记录:火山方舟上线 Doubao-Seed-2.1-Pro,用于开发者 API 调用和体验中心试用

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

CoWork-OS 使用 Grok 4.3 (high) 处理知识检索和问答

CoWork-OS · Grok 4.3 (high)

A
厂商:xAI / Grok 模型:Grok 4.3 (high) 来源平台:github 最后复核:2026-06-27T09:19:45Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CoWork-OS 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T09:19:45Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:CoWork-OS README 将项目定义为 local-first personal agentic OS / everything app,用于 coding、knowledge work、web design、automations 和 artifacts,并明确写明 Grok 支持 xAI API key 与 SuperGrok browser OAuth,grok-4.3 是默认 subscription model。

公开产物:公开产物是可访问的 CoWork-OS 开源仓库及配套 docs/providers.md;系统提供多 provider failover、prompt caching、Grok OAuth/API key 接入和本地优先 agent 工作区。

模型作用:Grok 4.3 在 CoWork-OS 中作为默认 Grok 订阅模型,承担个人 agentic OS 内的对话、编码/知识工作规划、自动化和 artifact 生成任务,并可通过 xAI 或 SuperGrok 凭据接入。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:README 明确模型绑定和产品任务,但未提供单个终端用户的运行截图;按公开开源产品集成案例采集。

原始记录:CoWork-OS 将 Grok 4.3 接入本地优先个人 Agentic OS

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

mLLMCelltype / cafferychen777 使用 Qwen3.6 Max Preview 处理文档理解和结构化处理

mLLMCelltype / cafferychen777 · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:github 最后复核:2026-06-27T10:10:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

mLLMCelltype / cafferychen777 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T10:10:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:mLLMCelltype is an open-source R/Python and web application for automated cell-type annotation in single-cell RNA sequencing. The Qwen provider module explicitly documents qwen3.6-max-preview as a supported Alibaba Qwen…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The tool returns per-cluster cell-type annotations, consensus predictions across multiple LLMs, uncertainty metrics, and documented reasoning that can be used inside Scanpy/Seurat workflows or the mllmcelltype.com web a…

模型作用:Qwen3.6 Max Preview 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3.6 Max Preview contributes one of the model opinions in the multi-LLM consensus process: it receives marker-gene and tissue/species prompts and returns candidate cell-type labels and reasoning that are combined wit…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The exact model is cited in provider code and the README's supported Alibaba Qwen model list. The public repository describes the artifact and task; individual user datasets are not included in the evidence.

原始记录:mLLMCelltype integrates Qwen3.6 Max Preview for consensus cell-type annotation of scRNA-seq marker genes

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

myuan19 / VoiceInput 使用 Qwen3.7 Max 处理多模态内容处理

myuan19 / VoiceInput · Qwen3.7 Max

A
厂商:Qwen / Alibaba 模型:Qwen3.7 Max 来源平台:github 最后复核:2026-06-27T09:24:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

myuan19 / VoiceInput 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T09:24:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:VoiceInput 是一个 Windows 托盘语音输入工具,用户录音后调用 DashScope ASR,并在配置中提供文本润色模型菜单;README 的配置表明确把 qwen3.7-max 列入默认启用的 enabled_polish_models,源码 config.py 也把 qwen3.7-max 和日期快照作为 Qwen 润色模型选项。

公开产物:公开仓库提供可下载 Releases、托盘录音、自动粘贴、历史记录和配置文件等完整桌面工具产物;用户可在配置中启用/选择 qwen3.7-max 对 ASR 文本进行润色后输出到当前光标位置。

模型作用:Qwen3.7 Max 在该工具中承担语音识别后文本整理/润色的 LLM 角色,把原始转写内容改写为更适合直接粘贴使用的文本;DashScope API Key 与模型菜单共同构成可核验的模型调用路径。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:个人开源桌面工具;证据能核验模型绑定、任务和公开产物,但没有公开生产用量或商业客户数据。

原始记录:myuan19 VoiceInput 用 Qwen3.7 Max 做语音转文字后的文本润色模型

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

GreyDGL / PentestGPT 使用 Qwen3.7 Max 处理智能体流程编排

GreyDGL / PentestGPT · Qwen3.7 Max

A
厂商:Qwen / Alibaba 模型:Qwen3.7 Max 来源平台:github 最后复核:2026-06-27T09:24:23Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

GreyDGL / PentestGPT 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T09:24:23Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:PentestGPT 是用于自动化渗透测试的开源 Agent/CLI;README 说明 legacy 版本运行多个协作 LLM 会话并支持 OpenAI、Anthropic、Gemini、DeepSeek、Qwen 等提供商,Supported LLM Providers 表明确将 Alibaba Qwen 的当前模型列为 qwen3.7-max,registry.py 也把 qwen3.7-max 注册为 Qwen flagship 模型。

公开产物:公开产物包括 PentestGPT CLI、官方演示、安装与运行命令、pentestgpt-legacy --list-models / --smoke-test 等操作入口;用户可把 QWEN_API_KEY 或 DASHSCOPE_API_KEY 配置给该工具,在渗透测试目标分析、推理、解析和攻击路径规划流程中调用 Qwen3.7 Max。

模型作用:Qwen3.7 Max 作为 PentestGPT 支持的 Qwen 旗舰模型,为渗透测试 Agent 的推理/解析 LLM 会话提供对话补全能力,参与目标信息分析、下一步测试动作建议和结果解释。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:安全工具存在双重用途风险;证据来自公开 README 与模型注册表,证明模型可被该工具作为工作流后端使用,但不代表官方客户背书或默认模型。

原始记录:PentestGPT Legacy 将 Qwen3.7 Max 接入自动化渗透测试 LLM 工作流

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Clerveu 使用 Qwen3.5 397B A17B 处理软件工程任务执行

Clerveu · Qwen3.5 397B A17B

A
厂商:Qwen / Alibaba 模型:Qwen3.5 397B A17B 来源平台:GitHub 最后复核:2026-06-27T09:36:14Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Clerveu 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:36:14Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:公开仓库 vidtest 提供一个视频转文本分析流水线:从 Source/ 选择视频,抽取/下载字幕,将视频切成 30 秒、720p、2fps 的片段,通过 OpenRouter 调用 Qwen3.5-397B-A17B vision model 为每个片段生成约 200 词视觉描述,再把字幕和视觉描述合并为带 [DIALOGUE] 与 [VISUAL DESCRIPTION] 的分析文件;配套 vlc_remote.py 支持 VLC 时间戳/截图交互。

公开产物:README 明确写明该项目使用 OpenRouter 的 Qwen3.5-397B-A17B vision model,并列出 7 步 pipeline、输出目录 descriptions/ 与 results/、OPENROUTER_MODEL = qwen/qwen3.5-397b-a17b;仓库中的 analyze.py 也公开包含 OPENROUTER_MODEL 常量、send_to_qwen() 调用和将模型描述写入结果文件的逻辑。

模型作用:Qwen3.5 397B A17B 是该流水线的核心视觉语言模型,负责读取视频片段与匹配字幕并生成结构化视觉描述;OpenRouter 提供 API 路由,项目脚本负责视频切片、字幕处理、结果合并和可选 Claude Code 后处理。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 仓库 README 与源码自述,未独立验证其实际运行规模或生产部署;但 exact model slug、使用者、任务流程和可访问 artifact 均明确。

原始记录:Clerveu 用 Qwen3.5-397B-A17B 构建视频转文本分析流水线

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

5dmgmt 使用 GPT-5.5 (xhigh) 处理软件工程任务执行

5dmgmt · GPT-5.5 (xhigh)

A
厂商:OpenAI 模型:GPT-5.5 (xhigh) 来源平台:GitHub 最后复核:2026-06-27T09:39:02Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

5dmgmt 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:39:02Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:5dmgmt 的 claude-codex-audit-toolkit 公开描述了一套 Claude Code + Codex CLI 审计循环:Claude Code 起草教材、审计 runbook 或实现代码后,由 Codex CLI 使用 GPT-5.5 xhigh 进行独立第三方审计,逐轮输出 finding、反映修复并用收敛判定决定继续、scope cut 或 ALL PASS。

公开产物:公开仓库提供 docs/01-08、examples、五月雨防止提示、环境 lint 清单、收敛判定准则和运行模板;overview 明确记录该方法已在 Workshop Course 1 runbook、SIFT Phase C runbook、video-subtitler 代码管线等 3 个内部案例中实走验证,分别达到 4 轮 scope cut、13 轮 ALL PASS、5 轮代码管线审计等结果。

模型作用:GPT-5.5 (xhigh) 作为 Codex CLI 的审计模型,负责从独立视角检查 Claude Code 产物的缺陷、环境陷阱、scope creep、收敛状态和代码/文档质量,并生成可逐轮处理的审计 findings 与结论。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 文档和仓库;案例记录了 5dmgmt 内部实走案例与工具包产物,但未公开完整逐轮原始 Codex 会话日志。

原始记录:5dmgmt 用 GPT-5.5 xhigh/Codex CLI 审计 Claude Code 产出的文档与代码库

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Magpie-Align 使用 Qwen2.5 72B 处理真实任务执行

Magpie-Align · Qwen2.5 72B

A
厂商:Qwen / Alibaba 模型:Qwen2.5 72B 来源平台:GitHub / Hugging Face 最后复核:2026-06-27T09:39:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Magpie-Align 公开的真实任务执行案例,来源为 公开代码库、Hugging Face 公开空间,复核于 2026-06-27T09:39:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Magpie-Align 使用 Qwen2.5 72B Instruct 作为生成模型,通过 Magpie 方法生成用于监督微调的原始 instruction-response conversations。

公开产物:公开发布 Magpie-Qwen2.5-Pro-1M-v0.1 数据集;项目导航页明确写明该数据集是 1M raw conversations built with Qwen2.5 72B Instruct,并提供 300K filtered 版本。

模型作用:Qwen2.5 72B Instruct 负责生成对齐数据中的用户指令和/或回答内容,是该 100 万条 SFT 对话数据的核心生成来源。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自 Magpie 官方仓库导航页和 Hugging Face 数据集页;这是研究/数据集产物,不是终端用户 SaaS 部署,但具有公开可核验产物。

原始记录:Magpie-Align 用 Qwen2.5 72B Instruct 生成 100 万条对齐训练对话数据

已有真实案例 真实任务执行公开代码库、Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

calderbuild 使用 Qwen2.5 72B 处理智能体流程编排

calderbuild · Qwen2.5 72B

A
厂商:Qwen / Alibaba 模型:Qwen2.5 72B 来源平台:GitHub / product demo 最后复核:2026-06-27T09:39:08Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

calderbuild 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T09:39:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:AI Smooth Talker 是一个面向对话场景的 AI 助手,用于在聊天、职场沟通等情境下生成多种高情商回复建议。

公开产物:项目 README 明确说明应用通过 SiliconFlow API 使用 Qwen/Qwen2.5-72B-Instruct 模型,并提供可访问的在线产品 Demo 页面。

模型作用:Qwen2.5 72B Instruct 负责根据用户输入生成自然、具有情绪智能的候选回复,是应用核心 AI 回复能力的生成模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:公开仓库 README 将模型绑定到 Qwen/Qwen2.5-72B-Instruct;Demo 页面可访问。仓库提示 API key 曾硬编码,未读取或记录任何密钥内容。

原始记录:calderbuild 用 Qwen2.5 72B Instruct 构建高情商回复生成助手 AI Smooth Talker

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

hanif017363 / St. Valentina AI Heal… 使用 DeepSeek R1 0528 处理医疗和生命科学分析

hanif017363 / St. Valentina AI Health Assistance · DeepSeek R1 0528

A
厂商:DeepSeek 模型:DeepSeek R1 0528 来源平台:github 最后复核:2026-06-27T09:37:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

hanif017363 / St. Valentina AI Health Assistance 公开的医疗与生命科学案例,来源为 公开代码库,复核于 2026-06-27T09:37:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Ai-Health-Assistance 是一个公开的 Web 健康聊天助手,页面标题和界面文案显示为 St. Valentina / Your Health Assistance,用户输入消息后由前端调用 OpenRouter Chat Completions API 生成回复。

公开产物:公开 GitHub 仓库提供可运行的 HTML/CSS/JavaScript 聊天界面;聊天机器人按系统提示以简短、礼貌的方式回答健康相关问题,并对非健康问题说明不属于其专长。

模型作用:script.js 在 aiChatRouter 中把请求模型固定为 deepseek/deepseek-r1-0528,并用系统提示约束其作为健康 chatbot 的回复风格和范围,因此 DeepSeek R1 0528 负责生成最终健康问答文本。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开应用源码;这是具体健康聊天应用而非 benchmark、教程或集合页。健康建议类应用存在医疗建议边界风险,适合在模型卡中注明不能替代专业医生。

原始记录:St. Valentina AI Health Assistance 使用 DeepSeek R1 0528 提供健康问答聊天

已有真实案例 医疗与生命科学公开代码库A 类可核验real_case auto_approved 进入模型卡精选

iluyobrainy / Deepdgm 使用 DeepSeek R1 0528 处理研究分析和报告生成

iluyobrainy / Deepdgm · DeepSeek R1 0528

A
厂商:DeepSeek 模型:DeepSeek R1 0528 来源平台:github 最后复核:2026-06-27T09:37:57Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

iluyobrainy / Deepdgm 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T09:37:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Deepdgm 是一个面向 secp256k1 椭圆曲线离散对数问题(ECDLP)的自治数学研究系统,README 说明其结合 Darwin Gödel Machine、自改进 AI 架构、DeepSeek R1 数学推理、SageMath 严格计算和自动验证。

公开产物:公开 GitHub 仓库提供数学研究 Agent 源码、自改进步骤和工具调用流程;系统目标是生成 ECDLP 研究策略、调用数学工具并保存自改进运行元数据和补丁草稿。

模型作用:llm_withtools.py 将 DEEPSEEK_MODEL 设为 deepseek-r1-0528,并在 chat_with_agent 中用该模型驱动带数学工具的研究对话;self_improve_step.py 也显式 create_client('deepseek-r1-0528') 来分析研究方法并生成改进建议。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开源码和 README;项目目标具有研究探索性质,不代表已解决 ECDLP,但模型被绑定到具体数学研究 Agent 的推理与自改进流程。

原始记录:Deepdgm 使用 DeepSeek R1 0528 构建 ECDLP 数学研究 Agent

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Asish-baidya29 / PaperTalk project 使用 PALM-2 处理代码审查和测试生成

Asish-baidya29 / PaperTalk project · PALM-2

A
厂商:Google / Gemini 模型:PALM-2 来源平台:github 最后复核:2026-06-27T09:34:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Asish-baidya29 / PaperTalk project 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T09:34:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:PaperTalk is a public Streamlit project for uploading PDFs, generating smart summaries, and chatting with document contents through an interactive retrieval interface.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository README describes an interactive PDF-based information-retrieval system that lets users upload a PDF, receive summaries, and ask questions about the document through a Streamlit UI.

模型作用:PALM-2 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:PaLM 2 is explicitly named as the language model used with FAISS and LangChain, providing the summarization and conversational question-answering layer over retrieved PDF content.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub README that directly names PaLM 2 and the concrete PDF chat task; small individual project/demo rather than a verified production customer deployment.

原始记录:PaperTalk used PaLM 2, FAISS and LangChain for PDF summarization and document chat

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Gitesh08 / Movie-Recommendation-Sys… 使用 PALM-2 处理软件工程任务执行

Gitesh08 / Movie-Recommendation-System project · PALM-2

A
厂商:Google / Gemini 模型:PALM-2 来源平台:github 最后复核:2026-06-27T09:34:49Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Gitesh08 / Movie-Recommendation-System project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:34:49Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The public Streamlit project takes a movie title, prompts a language model for related movie recommendations, and then enriches those generated titles through TMDB search results.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository code defines a recommendation flow that calls google.generativeai.generate_text with model 'models/text-bison-001', parses the returned newline-separated recommendations, and queries TMDB for movie metada…

模型作用:PALM-2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:PaLM 2's text-bison-001 model is the generative component that maps the user's input movie into a list of recommended titles before the application fetches external movie details.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is public source code with an explicit models/text-bison-001 binding and PALM_API_KEY configuration; project appears to be a small developer demo, not an enterprise deployment.

原始记录:A movie recommendation app used PaLM 2 text-bison-001 to generate similar-title suggestions

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

DAIR.AI / dair-ai 使用 MiniMax-M2.1 处理软件工程任务执行

DAIR.AI / dair-ai · MiniMax-M2.1

A
厂商:MiniMax 模型:MiniMax-M2.1 来源平台:github 最后复核:2026-06-27T09:46:43Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

DAIR.AI / dair-ai 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:46:43Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The public m2-deep-research repository describes a MiniMax-M2.1 Deep Research Agent powered by Minimax M2.1 with interleaved thinking, Exa neural search, and multi-agent orchestration. The README assigns MiniMax M2.1 to…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a runnable CLI research tool that decomposes a user query, searches the web with Exa, and generates comprehensive research reports with table of contents, key takeaways, executive summary, detail…

模型作用:MiniMax-M2.1 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax-M2.1 acts as the supervisor and synthesis model: it maintains reasoning state through interleaved thinking, coordinates the multi-step research workflow, and produces the final cited research report from retriev…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an open-source project README rather than a customer-story article; however it explicitly binds MiniMax M2.1 to concrete supervisor/synthesis roles in a public runnable artifact.

原始记录:DAIR.AI built a MiniMax-M2.1 deep research agent for multi-step web research reports

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Google Research / google-research-d… 使用 Gemini 1.0 Ultra 处理软件工程任务执行

Google Research / google-research-datasets · Gemini 1.0 Ultra

A
厂商:Google / Gemini 模型:Gemini 1.0 Ultra 来源平台:github 最后复核:2026-06-27T09:46:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Google Research / google-research-datasets 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:46:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Google Research published the Education Dialogue Dataset; its README states that the conversations were generated by prompting Gemini Ultra, with a teacher persona teaching a specified topic and a student persona respon…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public GitHub repository contains a released education-dialogue dataset with train and test JSON files, including multi-turn teacher/student conversations plus metadata such as topic, student preferences, teacher pr…

模型作用:Gemini 1.0 Ultra 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 1.0 Ultra, referenced by the project as Gemini Ultra, generated the core synthetic teacher-student dialogue content from the dataset prompts, making the model the production source for the released conversation a…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The README names Gemini Ultra rather than spelling out 'Gemini 1.0 Ultra'; because this public dataset was released in the Gemini 1.0 Ultra era and the task allows model-family binding, it is mapped to gemini-1-0-ultra.…

原始记录:Google Research used Gemini Ultra to generate the Education Dialogue Dataset

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

H2O.ai 使用 Llama 65B 处理研究分析和报告生成

H2O.ai · Llama 65B

A
厂商:Meta / Llama 模型:Llama 65B 来源平台:Hugging Face + GitHub 最后复核:2026-06-27T09:42:46Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

H2O.ai 公开的研究与报告生成案例,来源为 公开代码库、Hugging Face 公开空间,复核于 2026-06-27T09:42:46Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:H2O.ai 将 LLaMA 65B 作为 base model,使用 h2oai/openassistant_oasst1_h2ogpt_graded 数据集进行指令微调,发布 h2ogpt-research-oasst1-llama-65b,用于英文指令跟随、问答和聊天机器人场景;模型卡同时链接 h2oGPT GitHub 作为运行自有 chatbot 的代码入口。

公开产物:公开产物是 Hugging Face 上的 h2oai/h2ogpt-research-oasst1-llama-65b 模型仓库,模型卡明确写明这是 65B 参数 instruction-following LLM,列出 base model、微调数据集、训练日志、Transformers 调用示例和运行自有 chatbot 的 h2oGPT GitHub 链接。

模型作用:LLaMA 65B 提供 65B 参数基础语言能力;H2O.ai 在该底座上进行 OASST1/h2oGPT 数据指令微调,把基础模型转化为可用于聊天、问答和指令跟随的 h2oGPT 65B 研究模型。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:非商业许可/Meta LLaMA 1 权重访问限制仍适用;这是公开研究模型与聊天机器人产物,不代表商业客户部署。证据页明确绑定 base model 为 decapoda-research/llama-65b-hf,模型仓库公开可访问。

原始记录:H2O.ai 基于 LLaMA 65B 微调 h2oGPT OASST1 指令聊天模型

已有真实案例 研究与报告生成公开代码库、Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

豆包(ByteDance) 使用 Seed2.1 Pro 处理智能体流程编排

豆包(ByteDance) · Seed2.1 Pro

A
厂商:ByteDance Seed 模型:Seed2.1 Pro 来源平台:web_search 最后复核:2026-06-27T09:49:21Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 96/100

A 高可信 · 网页线索

原始证据1 个公开产物复核通过网页线索

豆包(ByteDance) 公开的智能体工作流案例,来源为 公开网页,复核于 2026-06-27T09:49:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:公开报道说明豆包专业版基于豆包 2.1 系列大模型,上线专属 2.1 Pro 旗舰模型与全新办公任务 Agent 模式;该模式面向复杂办公和生产力任务,支持操作本地电脑、使用浏览器、调用 Skills 与定时任务,并内置 Office 办公套件,可处理文档、表格、PPT、专业图片视频设计和生成可分享的应用网站。

公开产物:豆包官网提供公开产品入口;报道列举办公任务模式可整理本地资料、归类文件、处理文档、填写表格、跨应用协作、执行浏览器操作,并可搭建数据仪表盘、项目管理看板、活动报名系统、内容管理后台等带后端数据库和访问权限管理的在线应用。

模型作用:Seed2.1 Pro 在豆包专业版中作为旗舰模型支撑办公任务 Agent:理解用户工作目标,拆解多步骤任务,结合本地电脑、浏览器、文档/表格/PPT、Skills 与定时任务持续执行,并将需求转化为可交付的办公文件、视觉内容或在线应用产物。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开产品报道并可核验到豆包产品入口;这是豆包自身产品级部署案例,不是第三方客户故事或单个终端用户复盘。与已有 TRAE 案例不同,任务产物聚焦豆包专业版办公 Agent 和应用生成场景。

原始记录:豆包专业版接入 2.1 Pro,用办公任务 Agent 生成文档、表格、PPT 与可分享应用网站

已有真实案例 智能体工作流公开网页A 类可核验real_case auto_approved 进入模型卡精选

openfree 使用 Qwen3 235B 处理研究分析和报告生成

openfree · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:Hugging Face Spaces 最后复核:2026-06-27T09:49:42Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

openfree 公开的研究与报告生成案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T09:49:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:openfree published a Gradio application titled qwen3-235b-a22b Research for real-time deep research: it extracts search keywords from a user question with Qwen3-235B-A22B, searches the web through SerpHouse, and uses th…

公开产物:公开材料提供公开 Space、模型社区页面或源码,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Hugging Face Space and source code artifact implementing a web research assistant backed by accounts/fireworks/models/qwen3-235b-a22b.

模型作用:Qwen3 235B 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B is the core LLM invoked through Fireworks for keyword extraction and answer generation in the research workflow, with the Space metadata also binding the app to Qwen/Qwen3-235B-A22B.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public Space metadata and app.py source; the live hf.space runtime may require provider/API keys, so the durable artifact is the Hugging Face Space repository and source code.

原始记录:openfree built a Qwen3-235B-A22B real-time deep research Space

已有真实案例 研究与报告生成Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

kweheliye 使用 o3 处理软件工程任务执行

kweheliye · o3

A
厂商:OpenAI 模型:o3 来源平台:GitHub 最后复核:2026-06-27T09:42:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kweheliye 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:42:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Public FastAPI and LangChain project for processing energy deal confirmation memos: it accepts raw memo text, extracts structured deal fields, validates the result, and runs a counterparty credit-check tool before retur…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The repository provides a runnable server with /process-deal and /health endpoints plus sample curl usage; README describes a LangChain + OpenAI o3 agent, agent.py implements the two-turn extraction/tool-call/final-JSON…

模型作用:o3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:OpenAI o3 is the core reasoning model used through LangChain ChatOpenAI to parse unstructured deal memos, decide and execute the credit-check tool call, and synthesize the final structured response for downstream valida…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a small public GitHub repository rather than a customer story; however, the README, agent implementation, and config all bind the concrete application to OpenAI o3, and both repository and evidence URLs…

原始记录:Energetech deal processing agent uses OpenAI o3 to extract and validate energy trade memos

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ginipick 使用 Qwen3 235B 处理真实任务执行

ginipick · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:Hugging Face Spaces 最后复核:2026-06-27T09:49:42Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

ginipick 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T09:49:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ginipick published a Gradio Space that loads models/Qwen/Qwen3-235B-A22B through Hugging Face Inference Providers using the novita provider, giving signed-in users an interactive public demo for chatting with the model.

公开产物:公开材料提供公开 Space、模型社区页面或源码,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A reachable public Hugging Face Space and hf.space runtime exposing an interactive Qwen3-235B-A22B inference demo.

模型作用:Qwen3 235B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-235B-A22B is the explicitly loaded model in the Gradio app via gr.load('models/Qwen/Qwen3-235B-A22B', provider='novita'), providing the generated chat responses for the demo.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a community-built public demo rather than an enterprise customer story; use is still directly evidenced by source code and reachable artifact URLs.

原始记录:ginipick published a Hugging Face Inference Provider demo for Qwen3-235B-A22B

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

Harvey 使用 o1-preview 处理软件工程任务执行

Harvey · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:official_customer_story 最后复核:2026-06-27T09:42:50Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Harvey 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T09:42:50Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Harvey runs its professional-services legal AI platform on Azure AI infrastructure, including Azure OpenAI o1-preview, to help lawyers summarize and compare documents, perform due diligence, reference case law, analyze …

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The Microsoft customer story says Harvey is deployed across hundreds of law firms and legal teams, used by tens of thousands of lawyers, and that one corporate lawyer reported saving 10 hours of work per week; a Europea…

模型作用:o1-preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The story states Harvey's engineers primarily use Azure OpenAI Service o1-preview and o1-mini for reasoning and problem-solving tasks with increased focus, augmented with proprietary data for specific legal tasks and cu…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence explicitly names Azure OpenAI o1-preview, but Harvey also uses o1-mini and GPT-series models, so the public story does not isolate every product result to o1-preview alone.

原始记录:Harvey uses Azure OpenAI o1-preview in its legal AI platform for document analysis and due diligence

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

pinu 使用 Qwen2.5 Max 处理智能体流程编排

pinu · Qwen2.5 Max

A
厂商:Qwen / Alibaba 模型:Qwen2.5 Max 来源平台:huggingface_spaces 最后复核:2026-06-27T09:43:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

pinu 公开的智能体工作流案例,来源为 huggingface_spaces,复核于 2026-06-27T09:43:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The user published a Gradio web chatbot that lets visitors chat with Qwen2.5-Max, edit the system prompt, clear conversation state, and receive streamed model responses through Alibaba DashScope.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Hugging Face Space and live hf.space artifact titled Qwen2.5 Max Demo, with app.py exposing a Gradio chatbot UI and a DashScope Generation.call backend using model='qwen-max-0125'.

模型作用:Qwen2.5 Max 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:The Space source code labels the UI as Qwen2.5-Max/qwen2.5-max and calls DashScope Generation with the qwen-max-0125 model endpoint, so Qwen2.5-Max provides the assistant responses for the deployed chatbot demo.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Demo-scale individual Space rather than a production customer story; however the public source, live artifact, model endpoint, user, task, and output are explicit and reachable.

原始记录:pinu published a Hugging Face Gradio chatbot demo backed by Qwen2.5-Max

已有真实案例 智能体工作流huggingface_spacesA 类可核验real_case auto_approved 进入模型卡精选

willie-yao 使用 Qwen3 235B 处理真实任务执行

willie-yao · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:github 最后复核:2026-06-27T09:51:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

willie-yao 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T09:51:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:将 Cluster API Provider Azure 的 Prow 失败构建证据采集与摘要看板切换到私有集群内的 OpenAI-compatible Dynamo endpoint;README 明确 Demo 模型为 Qwen3-235B-A22B-FP8,AI_MODEL 设置为 Qwen/Qwen3-235B-A22B-FP8,并通过 self-hosted ARC runner 访问 qwen3-235b-a22b-disagg-frontend 服务。

公开产物:公开发布了 CAPZ Prow Dashboard 的 Qwen 版本页面,用同一套 prow jobs、prompt 和 evidence 生成真实 CAPZ 失败摘要;仓库说明其用于验证通用 AI path、私有模型集群内推理和不让数据离开集群的端到端 AI plumbing。

模型作用:Qwen3-235B-A22B-FP8 作为该 Demo 的核心故障摘要模型,接收 prow job、build、test 和 failure excerpt,生成已有真实案例到 dashboard 的失败解释与摘要;Hermes/qwen3 parser 负责结构化 tool calls 与 reasoning_content。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub README、workflow 和 GitHub Pages artifact;项目自称 demo/shadow variant,生产主看板仍使用 Claude,因此应标为公开工程 demo/验证案例而非最终生产替换。

原始记录:willie-yao 用 Qwen3-235B-A22B-FP8 驱动 CAPZ Prow 故障看板 Demo

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Altafk960 使用 Qwen3 235B 处理真实任务执行

Altafk960 · Qwen3 235B

A
厂商:Qwen / Alibaba 模型:Qwen3 235B 来源平台:github 最后复核:2026-06-27T09:51:06Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Altafk960 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T09:51:06Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Apache Spark 的 DAGScheduler、TaskSetManager 和 TaskSchedulerImpl 内嵌 LLMAdvisor,在每个 stage 边界调用 Cerebras OpenAI-compatible API 上的 qwen-3-235b-a22b-instruct,让模型根据分区、executor 状态、preferred locations 和近期 stage metrics 生成调度策略。

公开产物:仓库公开了 Spark 源码补丁、LLMAdvisor.scala、构建脚本和 POC;README 说明真实 Cerebras API 调用来自 Spark DAGScheduler 内部,8 个 workload stages 获得建议,约 140ms/次 LLM 调用,并在 API 失败或限流时回退到默认 Spark scheduling。

模型作用:Qwen3-235B 负责生成 JSON scheduling policy,包括 avoid_executors、prefer_executors、locality_override 和 speculation_hint;Spark 调度器随后用这些建议调整 executor offer 顺序、跳过不合适 executor、放宽/收紧 locality 并控制 speculative execution。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:公开证据为个人 GitHub 工程与 POC 说明,未见外部生产部署证明;但仓库包含可访问代码产物、明确模型绑定和具体 Spark 调度任务,不属于 benchmark-only 或教程-only。

原始记录:Altafk960 在 Apache Spark 调度器中接入 Cerebras Qwen3-235B 做实时任务放置决策

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Google 使用 Gemini 1.0 Ultra 处理智能体流程编排

Google · Gemini 1.0 Ultra

A
厂商:Google / Gemini 模型:Gemini 1.0 Ultra 来源平台:official_web 最后复核:2026-06-27T09:54:42Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Google 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T09:54:42Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Google 在 Gemini Advanced 中向用户提供 Gemini 1.0 Ultra,用于高复杂度对话协作,包括代码编写与调试、逻辑推理、遵循细致指令、创意项目协作、个性化学习辅导和长上下文对话。

公开产物:Gemini Advanced 作为公开可访问的 Gemini 产品体验上线;Google 官方说明该体验提供 Ultra 1.0,并称其在第三方盲评中相较领先替代聊天机器人更受偏好。

模型作用:Gemini 1.0 Ultra 是 Gemini Advanced 的核心高能力模型层,负责提供比 Pro 1.0 更强的复杂任务处理、长对话理解、编码、推理和创意协作能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是 Google 自家产品中的生产使用案例,证据来自官方发布页;artifact_url 为当前 Gemini 产品入口,页面内容可能已随新版本更新而不再突出 Ultra 1.0。

原始记录:Google 用 Gemini 1.0 Ultra 驱动 Gemini Advanced 付费 AI 助手

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

kernel-memory-dump 使用 o1-preview 处理软件工程任务执行

kernel-memory-dump · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:github 最后复核:2026-06-27T09:54:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

kernel-memory-dump 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:54:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The GitHub repository documents a desktop video-cropping application built with Electron and FFmpeg, and states it was 100% generated with a single ChatGPT o1-preview prompt. The app lets users load an MP4 file, visuall…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository contains the generated application and README usage instructions, including features for video preview, drag-based crop selection, progress tracking, crop scaling to the original resolution, instal…

模型作用:o1-preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:o1-preview is explicitly credited in the project title as the model used to generate the complete Electron/FFmpeg video-cropper application from one prompt, covering both the application implementation and documented us…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the creator's public GitHub repository rather than an independent customer story; the repository explicitly names ChatGPT o1-preview and exposes the artifact, but there is no separate production deploym…

原始记录:kernel-memory-dump generated an Electron FFmpeg video cropper with a single ChatGPT o1-preview prompt

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ZackTheDoer 使用 o1-preview 处理可玩交互原型构建

ZackTheDoer · o1-preview

A
厂商:OpenAI 模型:o1-preview 来源平台:github 最后复核:2026-06-27T09:54:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ZackTheDoer 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T09:54:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The repository records a prompt asking for a complete Space Invaders game with a neon aesthetic and heavy particle effects using Three.js, all in a single HTML file ready to copy and paste, and identifies the project as…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public artifact is a GitHub repository containing the generated single-file browser game and README instructions for opening index.html, moving the spaceship with arrow keys, shooting with spacebar, destroying alien…

模型作用:o1-preview 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:o1-preview is explicitly named by the repository and README as the model used to produce the complete game from one prompt, including the embedded JavaScript/HTML implementation and gameplay instructions.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public creator repository and not an enterprise deployment; it is still a concrete, reproducible artifact with exact o1-preview attribution and a clear generated software task.

原始记录:ZackTheDoer used o1-preview to create a single-file Three.js Space Invaders game

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Keakon 使用 DeepSeek-V2.5 处理软件工程任务执行

Keakon · DeepSeek-V2.5

A
厂商:DeepSeek 模型:DeepSeek-V2.5 来源平台:personal_blog 最后复核:2026-06-27T10:05:52Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Keakon 公开的代码代理与软件工程案例,来源为 博客记录,复核于 2026-06-27T10:05:52Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:作者在真实工程问题中使用 Cline 接入 DeepSeek-V2.5-1210,让模型把公司服务里同步调用 AWS Bedrock/Redrock Claude 的 Python 类改造成可异步调用的实现,并处理代理、凭证和调用代码问题。

公开产物:文章记录 DeepSeek-V2.5-1210 经过约半小时多轮修改后产出一个可以正常执行的版本,并展示了最终 Python 代码片段;作者同时指出仍存在额外代码、环境变量读取、usage 默认值和缺少自动刷新等问题。

模型作用:DeepSeek-V2.5-1210 作为 Cline 背后的代码生成与修改模型,读取上下文后生成 diff 和 Python 实现,帮助完成异步调用代码的主体改写。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:这是个人工程博客中的真实使用记录,不是官方客户故事;案例含横向评测语境,但有明确用户、任务、模型版本、产出代码和结果反馈。

原始记录:Keakon 用 Cline + DeepSeek-V2.5-1210 改造 AWS Bedrock Claude 调用为异步实现

已有真实案例 代码代理与软件工程博客记录A 类可核验real_case auto_approved 进入模型卡精选

junghyun-coding / Legis-Pilot 使用 o3 处理软件工程任务执行

junghyun-coding / Legis-Pilot · o3

A
厂商:OpenAI 模型:o3 来源平台:github 最后复核:2026-06-27T09:55:52Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

junghyun-coding / Legis-Pilot 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:55:52Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Legis-Pilot is a GovTech web application for the Korean Ministry of Government Legislation data contest. Citizens enter plain-language legislative proposals; the system compares them with official Korean legal-informati…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The product produces a validity judgment, recommended legislative form, a dashboard review, and a one-page report for ministry coordination, with a public live deployment and open GitHub repository documenting the flow.

模型作用:o3 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The README explicitly states that the four analytical axes are each handled by `o3` in parallel deep analysis before a chief-reviewer step synthesizes the final legislative feasibility report.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public project README and live demo; it documents a competition/product artifact rather than an independently published enterprise customer story. The repository binds the implementation to `o3`, bu…

原始记录:Legis-Pilot uses OpenAI o3 to assess Korean legislative proposals and generate one-page review reports

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

SillyTavern 使用 DeepSeek-V2.5 处理智能体流程编排

SillyTavern · DeepSeek-V2.5

A
厂商:DeepSeek 模型:DeepSeek-V2.5 来源平台:github 最后复核:2026-06-27T10:05:52Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

SillyTavern 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T10:05:52Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:SillyTavern 在其公开仓库的 release 分支内维护 DeepSeek-V2.5 instruct preset,为聊天/角色扮演应用配置 DeepSeek-V2.5 的 user、assistant、结束符等消息模板。

公开产物:公开产物是 default/content/presets/instruct/DeepSeek-V2.5.json,包含 input_sequence、output_sequence、output_suffix 等字段,使 SillyTavern 用户可在应用中按 DeepSeek-V2.5 的格式组织对话。

模型作用:DeepSeek-V2.5 的特殊对话标记和生成终止格式决定了该预设的模板内容,支撑 SillyTavern 将用户/助手消息正确包装后发送给模型进行角色对话生成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:该案例是开源产品集成配置而非单个终端用户故事;证据为可访问 GitHub 文件,明确绑定 DeepSeek-V2.5 并包含可复用 artifact。

原始记录:SillyTavern 为 DeepSeek-V2.5 提供内置 instruct 预设以支持角色对话

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

xvnpw / AI Security Analyzer 使用 Gemini 2.5 Pro (Mar) 处理研究分析和报告生成

xvnpw / AI Security Analyzer · Gemini 2.5 Pro (Mar)

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro (Mar) 来源平台:github 最后复核:2026-06-27T09:56:36Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

xvnpw / AI Security Analyzer 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T09:56:36Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AI Security Analyzer applies the Gemini 2.5 Pro March experimental model to an AI Nutrition-Pro application description/code input to generate security design, threat-modeling, attack-surface, attack-tree, and mitigatio…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public examples directory contains generated AI Nutrition-Pro security artifacts, including a detailed application threat model, attack tree, security design document, attack surface analysis, and mitigations; the G…

模型作用:Gemini 2.5 Pro (Mar) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro March experimental model performs the long-context security analysis and structured documentation generation, extracting assets, trust boundaries, data flows, attack paths, risks, and mitigation recommend…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub repository and generated example artifacts rather than a commercial customer story. The exact March experimental model ID is evidenced by the artifact filenames; the document body descri…

原始记录:AI Security Analyzer uses Gemini 2.5 Pro March experimental model to generate security documentation for AI Nutrition-Pro

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

TRAE / ByteDance 使用 Seed2.1 Turbo 处理软件工程任务执行

TRAE / ByteDance · Seed2.1 Turbo

A
厂商:ByteDance Seed 模型:Seed2.1 Turbo 来源平台:official_web 最后复核:2026-06-27T10:03:57Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

TRAE / ByteDance 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T10:03:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:TRAE IDE offers Doubao-Seed-2.1-Turbo as a built-in model option for developer workflows such as requirement analysis, feature implementation, bug fixing, environment setup, and result validation inside an AI coding env…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Developers can select Doubao-Seed-2.1-Turbo in TRAE IDE and use the model to produce and validate maintainable code changes across software-engineering tasks.

模型作用:Seed2.1 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Seed2.1 Turbo provides the coding-agent model capability behind the TRAE integration, contributing repository understanding, multi-step code generation, tool-use planning, and delivery of runnable engineering outputs.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from ByteDance Seed's official release article and the public TRAE product page; it is a first-party product integration rather than an independent third-party customer story.

原始记录:TRAE IDE exposes Doubao-Seed-2.1-Turbo for coding-agent workflows

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

Volcano Engine Ark / ByteDance 使用 Seed2.1 Turbo 处理多模态内容处理

Volcano Engine Ark / ByteDance · Seed2.1 Turbo

A
厂商:ByteDance Seed 模型:Seed2.1 Turbo 来源平台:official_web 最后复核:2026-06-27T10:03:57Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Volcano Engine Ark / ByteDance 公开的多模态生成与理解案例,来源为 官方页面,复核于 2026-06-27T10:03:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Volcano Engine Ark exposes Doubao-Seed-2.1-Turbo through its model-service platform so application developers can call the model from hosted API workflows and the Ark experience center.

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The model is available as a public platform option on Volcano Engine Ark, enabling developers to build applications that use Seed2.1 Turbo for agentic, coding, multimodal, and productivity tasks.

模型作用:Seed2.1 Turbo 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Seed2.1 Turbo supplies the underlying large-model inference capability for Ark-hosted applications, contributing lower-latency scalable generation and agent-oriented task execution for production API use.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a first-party availability/integration statement from the official Seed release and a public Ark product artifact; no independent customer deployment is claimed.

原始记录:Volcano Engine Ark makes Doubao-Seed-2.1-Turbo available for API applications

已有真实案例 多模态生成与理解官方页面A 类可核验real_case auto_approved 进入模型卡精选

Doubao / ByteDance 使用 Seed2.1 Turbo 处理软件工程任务执行

Doubao / ByteDance · Seed2.1 Turbo

A
厂商:ByteDance Seed 模型:Seed2.1 Turbo 来源平台:official_web 最后复核:2026-06-27T10:03:57Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

Doubao / ByteDance 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T10:03:57Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Doubao users can use the Seed2.1 family, including the Turbo variant, in office and productivity scenarios such as project planning, document processing, complex table analysis, report generation, and multi-step tool-as…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public Doubao chat product exposes an end-user workflow where Seed2.1 Turbo can be used to turn documents, tables, visual inputs, and user goals into usable office deliverables.

模型作用:Seed2.1 Turbo 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Seed2.1 Turbo contributes the general-agent and multimodal reasoning capabilities needed to understand task goals, process source materials, plan tool use, and generate structured deliverables for Doubao productivity wo…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is first-party Seed/Doubao product evidence; it establishes a concrete product integration but does not identify a separate external organization using the model.

原始记录:Doubao offers Seed2.1 Turbo for office-task mode and productivity workflows

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

swift-innovate 使用 Grok 4.20 0309 v2 处理代码审查和测试生成

swift-innovate · Grok 4.20 0309 v2

A
厂商:xAI / Grok 模型:Grok 4.20 0309 v2 来源平台:GitHub 最后复核:2026-06-27T10:28:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

swift-innovate 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T10:28:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:GOLEM 是一个基于 LangGraph 的自改进 Python Agent;项目文档说明它以 Grok 4.20 作为 primary brain,llm/grok.py 将 code_gen、code_review、self_improve、planning、reasoning 等任务路由到 grok-4.20-0309-reasoning,并配置自动 fallback。

公开产物:公开仓库包含可运行的 Agent 状态机、LLM 路由、预算、测试和自改进流程代码;Grok 路由模块通过 xAI OpenAI-compatible API 调用指定模型,为 Agent 的规划、代码生成、代码审查和自改进步骤提供输出。

模型作用:Grok 4.20 0309 Reasoning 被作为高强度推理与代码任务的主模型,负责 GOLEM Agent 的规划、推理、代码生成、代码审查和自改进决策。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据代码直接指定 grok-4.20-0309-reasoning;这是公开开源项目级使用证据,未独立验证外部生产流量或长期运行规模。

原始记录:swift-innovate 的 GOLEM 使用 Grok 4.20 0309 Reasoning 作为自改进 LangGraph Agent 主脑

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

panggi 使用 GPT-5.3 Codex (xhigh) 处理软件工程任务执行

panggi · GPT-5.3 Codex (xhigh)

A
厂商:OpenAI 模型:GPT-5.3 Codex (xhigh) 来源平台:github 最后复核:2026-06-27T10:04:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

panggi 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:04:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Generate book-style documentation/materials for Zircon Common Lisp, comparing outputs from Claude Code, Codex, and Kimi Code.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public GitHub repository containing generated Zircon Common Lisp book artifacts and a prompt file; the repository description explicitly says the Codex output used gpt-5.3-codex xhigh.

模型作用:GPT-5.3 Codex (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.3 Codex (xhigh) was one of the generation agents used to produce the Zircon Common Lisp book content, with the exact model variant named in the repository evidence.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub repository description and artifact repo; content should still be reviewed if model-card use requires separating Codex output from other model outputs.

原始记录:panggi used GPT-5.3 Codex xhigh to generate Zircon Common Lisp books

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

fvegiard 使用 Grok 4.20 0309 v2 处理软件工程任务执行

fvegiard · Grok 4.20 0309 v2

A
厂商:xAI / Grok 模型:Grok 4.20 0309 v2 来源平台:GitHub 最后复核:2026-06-27T10:05:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

fvegiard 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:05:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:omo-grok 是 Oh My OpenAgent + xAI Grok 的 OpenCode 双模型配置;Sisyphus-Grok Agent 定义明确使用 model: xai/grok-4.20-0309-reasoning,并将其作为默认编排器处理架构、规划、调试和深度思考。

公开产物:仓库公开了 OpenCode 项目配置、agent 定义和 oh-my-openagent 配置,可用于启动以 Grok 4.20 Reasoning 为核心的代码代理环境。

模型作用:Grok 4.20 0309 Reasoning 承担默认 orchestrator 和 deepsly-think 子代理职责,为代码项目中的规划、架构分析、调试和复杂推理提供高强度思考能力。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据文件中直接写明 xai/grok-4.20-0309-reasoning;属于公开配置产物,未声称生产部署规模。

原始记录:fvegiard 的 omo-grok 为 OpenCode 配置 Grok 4.20 0309 Reasoning 多 Agent 编程工作流

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Kenpal-ElenaTokumura 使用 GPT-5.3 Codex (xhigh) 处理可玩交互原型构建

Kenpal-ElenaTokumura · GPT-5.3 Codex (xhigh)

A
厂商:OpenAI 模型:GPT-5.3 Codex (xhigh) 来源平台:github 最后复核:2026-06-27T10:04:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Kenpal-ElenaTokumura 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T10:04:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build a Minesweeper game using Next.js, Bun, Tailwind CSS, lint/format/build scripts, and a documented local development workflow.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public repository for a working Next.js/Bun Minesweeper project with README instructions for installation, development, linting, formatting, building, and starting the app.

模型作用:GPT-5.3 Codex (xhigh) 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:The repository metadata states it was generated by GPT-5.3-Codex in GitHub Copilot Agent mode, indicating the model produced the application code and project scaffolding.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The public evidence names GPT-5.3-Codex but does not explicitly include the xhigh label; included because the task model is the GPT-5.3 Codex model entry and public model naming may omit reasoning-effort suffixes.

原始记录:Kenpal-ElenaTokumura generated a Next.js Minesweeper app with GPT-5.3 Codex in Copilot Agent mode

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

andreifie 使用 GPT-5.3 Codex (xhigh) 处理可玩交互原型构建

andreifie · GPT-5.3 Codex (xhigh)

A
厂商:OpenAI 模型:GPT-5.3 Codex (xhigh) 来源平台:github 最后复核:2026-06-27T10:04:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

andreifie 公开的游戏与交互原型案例,来源为 公开代码库,复核于 2026-06-27T10:04:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create an educational Phaser 3 game for engineering students to practice English grammar and vocabulary, with a Node.js/Express API, SQLite persistence, WebSocket social features, ranked play, lessons, friend challenges…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public full-stack game repository with client, server, database, real-time social features, API documentation, and local development/build commands.

模型作用:GPT-5.3 Codex (xhigh) 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:The repository description says the project was almost entirely vibecoded with GitHub Copilot plus gpt-5.3-codex, tying the model to implementation of the game and backend features.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The evidence names GPT-5.3-Codex without the xhigh suffix; included as a family-bound case for the GPT-5.3 Codex task model rather than as proof of a separately exposed xhigh runtime.

原始记录:andreifie built the FLYPER VIBE educational game largely with GitHub Copilot and GPT-5.3 Codex

已有真实案例 游戏与交互原型公开代码库A 类可核验real_case auto_approved 进入模型卡精选

x1colegal 使用 GPT-5.3 Codex (xhigh) 处理软件工程任务执行

x1colegal · GPT-5.3 Codex (xhigh)

A
厂商:OpenAI 模型:GPT-5.3 Codex (xhigh) 来源平台:github 最后复核:2026-06-27T10:04:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

x1colegal 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:04:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Prototype a peer-to-peer transport over plain HTTP/1.1 with peers serving object chunks, tracker-based peer discovery, HLS publish/ingest tooling, ffplay compatibility proxy, and a native Python/GStreamer player path.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Python repository implementing P2PHTTP, demo scripts for multi-peer operation, tracker/peer/client architecture documentation, and optional video playback tooling.

模型作用:GPT-5.3 Codex (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The repository metadata explicitly says the project was vibe-coded with GPT-5.3-Codex, indicating the model contributed to generating the implementation and documentation for the networking prototype.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The evidence names GPT-5.3-Codex without the xhigh suffix; included as a public artifact-bound use case for the GPT-5.3 Codex model family.

原始记录:x1colegal vibe-coded the P2PHTTP peer-to-peer HTTP transport with GPT-5.3 Codex

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

minhnh 使用 GPT-5.3 Codex (xhigh) 处理软件工程任务执行

minhnh · GPT-5.3 Codex (xhigh)

A
厂商:OpenAI 模型:GPT-5.3 Codex (xhigh) 来源平台:github 最后复核:2026-06-27T10:04:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

minhnh 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:04:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Create a Neovim plugin for robbdd/TextX languages, including filetype detection, syntax highlighting fallback, optional LSP diagnostics and hover documentation via a bundled Python language server, setup configuration, …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Neovim plugin repository with Lua/Python plugin code, README setup instructions, configuration examples, and Python/Neovim smoke test commands.

模型作用:GPT-5.3 Codex (xhigh) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:The repository metadata states the Neovim plugin was created using GPT 5.3-Codex, tying the model to the generated developer-tooling implementation.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The evidence names GPT 5.3-Codex without the xhigh suffix; included as a model-family-bound artifact because public GitHub descriptions commonly omit the reasoning-effort variant.

原始记录:minhnh created a Neovim plugin for robbdd/TextX languages using GPT-5.3 Codex

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

khalimanja 使用 Grok 4.20 0309 v2 处理智能体流程编排

khalimanja · Grok 4.20 0309 v2

A
厂商:xAI / Grok 模型:Grok 4.20 0309 v2 来源平台:GitHub 最后复核:2026-06-27T10:05:59Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

khalimanja 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T10:05:59Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:omni-bot 是一个 Flask Webhook Bot:接收 lead reply 后调用 xAI SDK 的 chat.create,model 设置为 grok-4.20-reasoning,并用销售 closer 系统提示生成回复。

公开产物:公开代码产物实现了 /webhook 接口、后台运行循环、Grok 调用和回复打印逻辑,可作为销售线索自动回复 Bot 的最小可运行服务。

模型作用:Grok 4.20 Reasoning 根据潜在客户回复生成直接、面向成交的销售响应,承担对话理解和自动文案生成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:代码明确指定 grok-4.20-reasoning;仓库是最小实现,发送回 Instantly 的集成仍标注 TODO,因此产物成熟度有限但证据可核验。

原始记录:khalimanja 的 omni-bot 使用 Grok 4.20 Reasoning 自动回复销售线索 Webhook

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

IndyDevDan / disler 使用 Gemini 2.5 Pro (Mar) 处理软件工程任务执行

IndyDevDan / disler · Gemini 2.5 Pro (Mar)

A
厂商:Google / Gemini 模型:Gemini 2.5 Pro (Mar) 来源平台:github 最后复核:2026-06-27T09:59:18Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

IndyDevDan / disler 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T09:59:18Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build and publish an MCP server that lets Claude Code or other MCP clients offload AI coding tasks to Aider; the README documents an Aider AI Code tool for refactoring and code-generation requests and configures Gemini …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public GitHub repository and installable Python project named aider-mcp-server, with README instructions for adding the MCP server to Claude Code using gemini/gemini-2.5-pro-exp-03-25 and a project script entry point …

模型作用:Gemini 2.5 Pro (Mar) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Gemini 2.5 Pro Experimental 03-25 is the default model for Aider's code-generation path in the server and is explicitly named in the README as the model used for tests and as the primary model default for generating cod…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is project-maintainer documentation rather than an external customer story; it is still a concrete public artifact with reachable repo, exact model string, user/maintainer identity, task and MCP coding workflow…

原始记录:Aider MCP Server defaults its AI coding tool to Gemini 2.5 Pro Experimental 03-25

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ruvnet / RuFlo 使用 Qwen3.6 Max Preview 处理文档理解和结构化处理

ruvnet / RuFlo · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:github 最后复核:2026-06-27T10:10:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ruvnet / RuFlo 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T10:10:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:RuVocal is a SvelteKit web UI for RuFlo that lets users chat with Qwen, Claude, Gemini, or OpenAI while invoking roughly 210 MCP tools for agent orchestration, persistent memory, swarm coordination, code review, GitHub …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public hosted demo at flo.ruv.io and a self-hostable Docker/Cloud Run deployment provide multi-model chat, parallel MCP tool execution, memory-backed conversations, and a tool gallery for end users.

模型作用:Qwen3.6 Max Preview 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3.6 Max Preview is the documented default/curated model route for generating chat responses and tool-calling decisions in the RuVocal interface, turning user requests and tool results into workflow automation action…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the public repository README and hosted product URL; live inference routing was not authenticated or instrumented, but the exact OpenRouter model id qwen/qwen3.6-max-preview is explicitly documented in …

原始记录:RuVocal hosted MCP chat uses Qwen 3.6 Max Preview as a curated/default model for tool-using workflow automation

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LeaderOnePro / DeepDrone 使用 Qwen3.6 Max Preview 处理真实任务执行

LeaderOnePro / DeepDrone · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:github 最后复核:2026-06-27T10:10:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LeaderOnePro / DeepDrone 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T10:10:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:DeepDrone is an AI-powered drone control system with terminal and web interfaces. Its interactive setup lists qwen3.6-max-preview under the Qwen provider via DashScope, and the README describes controlling drones with n…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project provides a runnable drone simulator, real DroneKit integration, and a web/CLI interface that converts conversational instructions into drone-control actions with telemetry and safety controls.

模型作用:Qwen3.6 Max Preview 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:When selected in the Qwen provider configuration, Qwen3.6 Max Preview acts as the language-planning layer that interprets user flight commands and helps produce structured drone-control intents/actions for the simulator…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence binds the exact model in the setup code and the artifact describes the drone-control application. The repository demonstrates selectable model support rather than a third-party deployment report.

原始记录:DeepDrone exposes Qwen3.6 Max Preview for natural-language drone control and simulator/real-flight command generation

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

qzw881130 / AI-NovelFlow 使用 Qwen3.6 Max Preview 处理文档理解和结构化处理

qzw881130 / AI-NovelFlow · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:github 最后复核:2026-06-27T10:10:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

qzw881130 / AI-NovelFlow 公开的文档理解与处理案例,来源为 公开代码库,复核于 2026-06-27T10:10:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕文档理解和结构化处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:AI-NovelFlow is an AI-driven platform that converts novels into videos. The public project documents workflow steps including novel import, AI character extraction, scene extraction, prop extraction, storyboard splittin…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The artifact is a full-stack FastAPI/React/ComfyUI application with screenshots and public demo videos for generating character libraries, scene references, storyboards, audio, and composed video from novel text.

模型作用:Qwen3.6 Max Preview 在该案例中承担文档理解和结构化处理相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3.6 Max Preview is available as a configured LLM option for the text-understanding stages of the pipeline, helping parse long-form novel content into structured characters, scenes, props, and storyboard prompts cons…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The exact model is present in public settings/configuration code and the repository documents the production workflow. Evidence does not prove it is the default model for every deployment, only that it is a supported re…

原始记录:AI-NovelFlow adds Qwen3.6 Max Preview as an Aliyun Bailian model for AI-assisted novel-to-video production

已有真实案例 文档理解与处理公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CanhuiLiPhy / 明史阅读器 使用 Qwen3.6 Max Preview 处理软件工程任务执行

CanhuiLiPhy / 明史阅读器 · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:github 最后复核:2026-06-27T10:10:19Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CanhuiLiPhy / 明史阅读器 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:10:19Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:明史阅读器 is a local Ming-dynasty history reading tool with 20+ historical texts, timelines, place names, offices, notes/bookmarks, and LLM-assisted reading. Its default model options include qwen3.6-max-preview through the DashScope/OpenAI-compatible configuration.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The application provides downloadable/source-build reading software with AI operations such as translation, pronunciation, explanation, encyclopedia lookup, cross-book historical-material comparison, AI punctuation/segm…

模型作用:Qwen3.6 Max Preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3.6 Max Preview is one of the configured large-model options that can power the reader's AI assistance over classical Chinese passages, producing explanations, translations, comparisons, and chat responses grounded …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The exact model is bound in the public default configuration and the README describes the AI reading artifact. The evidence is product-source documentation, not a customer story about one specific end-user session.

原始记录:明史阅读器 includes Qwen3.6 Max Preview for classical Chinese history reading assistance

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

abe238 / AI PM Resume Analyzer 使用 Claude 4.5 Sonnet 处理软件工程任务执行

abe238 / AI PM Resume Analyzer · Claude 4.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 4.5 Sonnet 来源平台:GitHub 最后复核:2026-06-27T10:12:11Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

abe238 / AI PM Resume Analyzer 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:12:11Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Analyze Applied AI PM resumes from PDF/DOC/DOCX files against a 2026 six-pillar framework covering technical skills, product thinking, AI/ML knowledge, communication, strategic thinking, and execution; users can run the…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The tool generates structured resume evaluations with 0-10 pillar scores, evidence quotes, total score out of 60, decision labels such as Strong Screen, Markdown/HTML/JSON reports, and optional deep-analysis consensus o…

模型作用:Claude 4.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Sonnet 4.5 is implemented as the default Anthropic model in the analyzer and is called to read resume text, optionally evaluate visual design from a resume image, apply the six-pillar rubric, return JSON scoring,…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public open-source repository and README verify the task, artifact, and exact Anthropic model id; this is an independent tool rather than an Anthropic customer story, and users must supply their own Anthropic API key wi…

原始记录:AI PM Resume Analyzer uses Claude Sonnet 4.5 to score resumes against a six-pillar AI PM hiring framework

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

rhadiwib 使用 MiniMax M1 80k 处理真实任务执行

rhadiwib · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:Hugging Face Spaces 最后复核:2026-06-27T10:16:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

rhadiwib 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T10:16:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上搭建 MiniMaxAI/MiniMax-M1-80k 的 Gradio 在线推理页面,登录 Hugging Face 后通过 novita provider 调用模型做文本生成/聊天体验。

公开产物:公开 Space 页面和源码均可访问;app.py 明确写有 This Space showcases the MiniMaxAI/MiniMax-M1-80k model,并通过 gr.load("models/MiniMaxAI/MiniMax-M1-80k", accept_token=button, provider="novita") 载入模型。

模型作用:MiniMax M1 80k 是该 Space 对外展示和执行文本生成/聊天响应的核心模型,负责把用户输入转换为模型回复。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开 Hugging Face 推理 Demo,不证明生产客户落地;但使用者、公开产物、源码证据和 exact model MiniMaxAI/MiniMax-M1-80k 绑定清晰,URL 已核验 200。

原始记录:rhadiwib 发布 MiniMax-M1-80k Hugging Face Gradio 推理 Demo

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

lindembergue 使用 MiniMax M1 80k 处理真实任务执行

lindembergue · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:Hugging Face Spaces 最后复核:2026-06-27T10:16:29Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 98/100

A+ 完整链路 · 公开产物空间

原始证据1 个公开产物复核通过公开产物空间

lindembergue 公开的真实任务执行案例,来源为 Hugging Face 公开空间,复核于 2026-06-27T10:16:29Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:在 Hugging Face Spaces 上发布一个用于测试 MiniMaxAI/MiniMax-M1-80k 的 Gradio 推理应用,README 标注 short_description 为 teste,并配置 Hugging Face OAuth inference-api 权限。

公开产物:公开 Space 页面可访问;app.py 明确声明该 Space showcases the MiniMaxAI/MiniMax-M1-80k model, served by the novita API,并调用 gr.load 载入 models/MiniMaxAI/MiniMax-M1-80k。

模型作用:MiniMax M1 80k 是该测试/演示 Space 的唯一载入模型,承担在线推理和文本输出生成。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据为公开测试型 Demo,业务场景较轻,不作为生产成效证明;但 exact model、使用者、任务、源码证据和公开 artifact 均可核验,URL 已核验 200。

原始记录:lindembergue 发布 MiniMax-M1-80k Hugging Face 推理测试 Space

已有真实案例 真实任务执行Hugging Face 公开空间A 类可核验real_case auto_approved 进入模型卡精选

Abe Diaz / abe238 使用 Claude 4.5 Sonnet 处理软件工程任务执行

Abe Diaz / abe238 · Claude 4.5 Sonnet

A
厂商:Anthropic / Claude 模型:Claude 4.5 Sonnet 来源平台:github 最后复核:2026-06-27T10:13:27Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Abe Diaz / abe238 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:13:27Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an open-source Applied AI Product Manager resume analyzer for job seekers and hiring teams, using a six-pillar AI PM evaluation framework to score PDF, DOC, and DOCX resumes and produce structured hiring-readiness…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public project publishes a runnable analyzer and product site; its README states it generates per-candidate Markdown, HTML, and JSON reports with 0-10 scoring by pillar, evidence, level assessment, and model-selecta…

模型作用:Claude 4.5 Sonnet 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude Sonnet 4.5 is explicitly listed as an Anthropic option and default Anthropic model for the analyzer; the CLI model table binds claude-sonnet-4-5-20250929 to complex resume-analysis work and the README documents r…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:The tool is multi-provider and defaults to OpenAI unless configured; accepted because the README and CLI code explicitly bind the Anthropic pathway to claude-sonnet-4-5-20250929 and provide a public artifact/product for…

原始记录:Abe Diaz built an AI PM resume analyzer that can use Claude Sonnet 4.5 for six-pillar hiring evaluations

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

teremterem 使用 GPT-5.1 (high) 处理软件工程任务执行

teremterem · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:18:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

teremterem 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:18:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开源项目提供一个本地 LiteLLM 代理,让 Anthropic Claude Code CLI 可以调用 OpenAI 模型;配置模板将 Claude Opus 默认重映射为 gpt-5.1-reason-high,并提供 Docker/uv 运行脚本和 Claude Code 连接方式。

公开产物:公开 GitHub 仓库包含可运行代理配置、环境变量模板和运行脚本,用户可通过 ANTHROPIC_BASE_URL 指向本地代理,在 Claude Code CLI 中使用映射后的 OpenAI GPT-5.1 high reasoning 模型。

模型作用:GPT-5.1 (high) 作为 Claude Opus 路由目标,为 Claude Code CLI 会话提供高推理强度的代码生成、编辑和工具调用能力;仓库还针对 GPT-5 medium/high reasoning 可能并行工具调用的问题提供单工具调用约束配置。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开仓库配置模板和 README;这是开源工具/集成案例,不是第三方客户故事。

原始记录:teremterem 将 Claude Code CLI 通过 LiteLLM 代理映射到 GPT-5.1 high reasoning

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Tinghecui 使用 GPT-5.1 (high) 处理软件工程任务执行

Tinghecui · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:18:17Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Tinghecui 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:18:17Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:开源项目 gpt-cc 基于 Claude Code/OpenAI 代理方案,提供 LiteLLM 配置与运行脚本;环境变量模板明确把 Claude Opus 重映射到 gpt-5.1-reason-high,用于在 Claude Code CLI 中走 OpenAI 模型。

公开产物:公开仓库给出了安装、配置、uv/Docker 启动以及通过 ANTHROPIC_BASE_URL 连接 Claude Code 的完整产物,形成可复用的本地开发代理工具。

模型作用:GPT-5.1 (high) 是该代理配置中面向高推理强度编码会话的目标模型,负责承接 Claude Code 中原本发往 Opus 的复杂代码理解、生成和编辑请求。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开仓库配置模板;该仓库说明其基于 teremterem/claude-code-gpt-5-codex,属于衍生开源集成案例。

原始记录:Tinghecui/gpt-cc 用 GPT-5.1 high reasoning 作为 Claude Code 的 OpenAI 后端

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

farion1231 / CC Switch 使用 Seed-2.1-Pro-Preview 处理软件工程任务执行

farion1231 / CC Switch · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:github 最后复核:2026-06-27T10:18:04Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

farion1231 / CC Switch 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:18:04Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:CC Switch 是面向 Claude Code、Claude Desktop、Codex、OpenCode、OpenClaw、Hermes 等 AI 编程工具的跨平台配置管理器。其公开 CHANGELOG 在 3.16.4 条目中明确记录:DouBaoSeed preset now targets doubao-seed-2-1-pro,替换 doubao-seed-2-0-code-preview-latest,并在 claude、claude-desktop、codex、opencode、openclaw、hermes 六个客户端中统一更新显示名和成本配置。

公开产物:公开仓库提供可下载/可构建的 CC Switch 桌面应用;v3.16.4 之后用户可以在六类客户端配置中直接选择 Doubao Seed 2.1 Pro 预设,用于切换 API provider、写入客户端配置、统计模型用量和按新价格归因成本。

模型作用:Doubao Seed 2.1 Pro 在该案例中作为 CC Switch 的内置 AI 编程模型预设,承担 Claude/Codex/OpenCode/OpenClaw/Hermes 等客户端背后的代码生成、Agent 编程和长任务执行模型角色;项目围绕该模型更新了模型 ID、展示名、价格和多客户端配置映射,使真实用户能把同一模型接入多个本地 AI coding 工具。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 仓库的 CHANGELOG 和可访问仓库产物;这是第三方开发者工具的模型接入案例,不是最终企业业务部署。模型 ID 在上游以 doubao-seed-2-1-pro 形式出现,与本任务的 Seed-2.1-Pro-Preview/正式 Pro 家族绑定;不使用 benchmark 或发布综述作为案例。

原始记录:CC Switch 将 Doubao Seed 2.1 Pro 设为六类 AI 编程客户端的内置预设

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Scira / zaidmukaddam 使用 MiniMax M1 80k 处理研究分析和报告生成

Scira / zaidmukaddam · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:GitHub 最后复核:2026-06-27T10:17:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Scira / zaidmukaddam 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T10:17:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Scira is an agentic research platform where users ask questions, upload PDFs, or paste URLs; the system plans subtasks, retrieves live sources, and produces cited answers. Its provider registry exposes a `scira-minimax`…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public research product and open-source app can route research-answer generation through MiniMax M1 80K and return grounded answers with inline citations for user queries.

模型作用:MiniMax M1 80k 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M1 80K supplies the reasoning/chat model behind Scira's MiniMax option, contributing long-context reasoning and answer synthesis in the agentic research workflow.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public production app plus source-level model integration; it does not expose private usage volume or telemetry.

原始记录:Scira uses MiniMax M1 80K as a selectable model for agentic cited research

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

OpenBMB / PilotDeck 使用 GPT-5.1 (high) 处理智能体流程编排

OpenBMB / PilotDeck · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:17:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

OpenBMB / PilotDeck 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T10:17:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:PilotDeck is a task-oriented AI agent productivity platform with workspaces and a live demo; its model selector includes the concrete option value gpt-5.1-high labelled GPT-5.1 High for agent tasks.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public open-source AI agent productivity platform and hosted demo where users can select GPT-5.1 High among supported models for workspace-based task execution.

模型作用:GPT-5.1 (high) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 High is provided as one of the reasoning-capable model choices powering PilotDeck's task agents, contributing model reasoning and generation during workspace task execution.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public product repository and code-level model integration; it does not disclose individual end-user transcripts or production volume.

原始记录:OpenBMB PilotDeck exposes GPT-5.1 High for task-oriented AI agent workspaces

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

CherryHQ / Cherry Studio 使用 MiniMax M1 80k 处理知识检索和问答

CherryHQ / Cherry Studio · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:GitHub 最后复核:2026-06-27T10:17:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

CherryHQ / Cherry Studio 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T10:17:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Cherry Studio App is the official mobile version of Cherry Studio for iOS and Android, providing multi-LLM conversations, assistants, history search, and data migration. Its default model catalog includes `minimaxai/min…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Mobile users of Cherry Studio App can select MiniMax M1 80K through the app's model catalog for assistant/chat interactions alongside other supported LLMs.

模型作用:MiniMax M1 80k 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M1 80K contributes the chat and reasoning backend option used by Cherry Studio's multi-model assistant experience on mobile.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is an application model catalog and public mobile app repository; no customer-specific deployment metrics are published.

原始记录:Cherry Studio App lists MiniMax M1 80K for mobile multi-LLM chat

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Paperclip AI 使用 GPT-5.1 (high) 处理研究分析和报告生成

Paperclip AI · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:17:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Paperclip AI 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T10:17:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Paperclip is an open-source app for orchestrating teams of AI agents at work; its Cursor local adapter includes gpt-5.1-high in the supported model list used to run agents from the dashboard.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public Node.js/React agent-management application that can run business-oriented agent workflows with gpt-5.1-high available as a concrete model option.

模型作用:GPT-5.1 (high) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 High contributes the LLM reasoning/generation layer for selected Paperclip agents when users configure that model for agent execution.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence verifies product integration and public artifact, but not a named customer deployment or metrics.

原始记录:Paperclip integrates gpt-5.1-high for managing teams of work agents

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

serithemage / Solar Code project 使用 Solar Pro 2 处理软件工程任务执行

serithemage / Solar Code project · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:github 最后复核:2026-06-27T10:12:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

serithemage / Solar Code project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:12:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The Solar Code project adapts a Gemini CLI-style command-line agent into a developer workflow tool for Korean developers and organizations, letting users query and edit codebases, generate applications, automate operati…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public repository provides an installable Node.js CLI, documentation, screenshots, configuration flow, and source code. Its model constants set DEFAULT_SOLAR_MODEL to solar-pro2 and the README states the tool is pow…

模型作用:Solar Pro 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 2 is the primary LLM behind the coding assistant: it supplies the language understanding and generation used for codebase Q&A, edits, application generation, PR/issue workflow assistance, and Korean/English de…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public GitHub engineering artifact rather than a named enterprise customer story; accepted because it is a concrete runnable product/tool repo with explicit Solar Pro2 binding and reachable implementation files, not a t…

原始记录:Solar Code builds a Solar Pro 2 powered command-line coding workflow assistant

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

chatless / kamjin3086 使用 MiniMax M1 80k 处理知识检索和问答

chatless / kamjin3086 · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:GitHub 最后复核:2026-06-27T10:17:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

chatless / kamjin3086 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T10:17:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:chatless is a Tauri and Next.js local-first desktop AI chat client supporting cloud and local models, document parsing, vision, local RAG knowledge bases, MCP tools, prompt management, and WebDAV prompt sync. Its static…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Desktop users can configure the chatless client to use MiniMax M1 80K for chat and knowledge-base/RAG workflows while keeping application data local.

模型作用:MiniMax M1 80k 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M1 80K serves as one of the cloud reasoning/chat model options available to generate answers, analyze documents, and support RAG-assisted conversations in the client.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is public source and a released desktop client repository; exact end-user usage counts are not published.

原始记录:chatless adds MiniMax M1 80K to a local-first desktop AI client

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

BloopAI / Vibe Kanban 使用 GPT-5.1 (high) 处理软件工程任务执行

BloopAI / Vibe Kanban · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:17:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

BloopAI / Vibe Kanban 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:17:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Vibe Kanban lets engineers plan work in kanban issues and run coding agents in isolated workspaces; its Cursor executor resolves base model gpt-5.1 with high/default reasoning to gpt-5.1-high and exposes high as a reaso…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public coding-agent management tool where GPT-5.1 High can be selected to execute software engineering tasks in agent workspaces with branches, terminals and dev-server previews.

模型作用:GPT-5.1 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 High supplies the high-reasoning model endpoint used by the Cursor executor for coding-agent task execution and iterative code changes.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Repository is a real public product; the specific evidence is implementation-level support rather than a published case study with quantitative outcomes.

原始记录:Bloop Vibe Kanban maps GPT-5.1 with high reasoning to gpt-5.1-high for coding-agent workspaces

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ho323 / MixUP Upstage Promptathon p… 使用 Solar Pro 2 处理软件工程任务执行

ho323 / MixUP Upstage Promptathon participant · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:GitHub 最后复核:2026-06-27T10:25:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

ho323 / MixUP Upstage Promptathon participant 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:25:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Convert Korean sentences containing archaic Korean, Hanja, classical Chinese, and mixed-language expressions into modern Korean news-style sentences for a MixUP Upstage Promptathon task.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository containing a Solar Pro2 prompt-engineering pipeline, CSV batch generation workflow, prompt configuration, and reported scores of 0.8360 on the public leaderboard and 0.8140 final score.

模型作用:Solar Pro 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro2 is the core generation model in a three-turn process: meaning-preserving modernization, fluency/news-style refinement, and verification/correction to reduce semantic drift.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public competition project README and code artifact; result scores are self-reported in the repository README and were not independently re-run.

原始记录:MixUP Upstage project uses Solar Pro2 for Korean historical-text modernization

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

joey-zhou / Xiaozhi ESP32 Server Ja… 使用 MiniMax M1 80k 处理多模态内容处理

joey-zhou / Xiaozhi ESP32 Server Java · MiniMax M1 80k

A
厂商:MiniMax 模型:MiniMax M1 80k 来源平台:GitHub 最后复核:2026-06-27T10:17:54Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

joey-zhou / Xiaozhi ESP32 Server Java 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T10:17:54Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Xiaozhi ESP32 Server Java is a Java backend and management platform for ESP32 smart-hardware dialogue, with WebSocket/MQTT communication, STT/TTS, IoT control, and multi-AI-platform integration. Its web LLM factory conf…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The smart-device backend can offer MiniMax M1 80K as a configurable chat model for voice assistant and IoT-control interactions from ESP32 devices.

模型作用:MiniMax M1 80k 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:MiniMax M1 80K contributes long-context chat/reasoning capability within the server's multi-provider LLM layer, enabling natural-language dialogue and tool-capable assistant responses.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public server repository and configuration artifact; live deployment volume is not disclosed.

原始记录:Xiaozhi ESP32 Server Java exposes MiniMax M1 80K for smart-hardware dialogue

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Terragon Labs 使用 GPT-5.1 (high) 处理软件工程任务执行

Terragon Labs · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:17:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Terragon Labs 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:17:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Terragon is a cloud orchestrator for delegating development work to coding agents in isolated sandboxes with GitHub PR workflows; its agent model list includes gpt-5.1-high for latest/v2 agent configurations.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:An open-source snapshot of a cloud coding-agent product that can delegate repository tasks, checkpoint changes and create pull requests using GPT-5.1 High as a supported Codex/OpenAI model option.

模型作用:GPT-5.1 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 High acts as the selected reasoning/generation model for Terragon coding agents when users choose that option for cloud task execution.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public snapshot after product shutdown; evidence verifies implementation and task flow, but not current hosted availability or customer metrics.

原始记录:Terragon cloud coding-agent orchestrator supports GPT-5.1 High for delegated development tasks

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

code-yeongyu / oh-my-openagent 使用 GPT-5.1 (high) 处理软件工程任务执行

code-yeongyu / oh-my-openagent · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:GitHub 最后复核:2026-06-27T10:17:31Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

code-yeongyu / oh-my-openagent 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:17:31Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:oh-my-openagent is a multi-harness agent OS for OpenCode, Codex, Pi and related coding agents; its think-mode switcher maps the GPT-5.1 model key gpt-5-1 directly to gpt-5-1-high for high-reasoning coding-agent runs.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:A public coding-agent harness where users can invoke the high-reasoning GPT-5.1 variant through the think-mode model switcher while running OpenCode/Codex-style development workflows.

模型作用:GPT-5.1 (high) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 High is the target high-reasoning model selected by the harness when users switch GPT-5.1 into think mode, contributing deeper reasoning for complex codebase tasks.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public open-source harness integration with exact gpt-5-1-high mapping; it does not include private runtime logs or outcome metrics.

原始记录:oh-my-openagent maps GPT-5.1 to gpt-5-1-high in its think-mode coding-agent harness

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

luwill / Claude-Code-Model-Router 使用 Seed-2.1-Pro-Preview 处理软件工程任务执行

luwill / Claude-Code-Model-Router · Seed-2.1-Pro-Preview

A
厂商:ByteDance Seed 模型:Seed-2.1-Pro-Preview 来源平台:github 最后复核:2026-06-27T10:12:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

luwill / Claude-Code-Model-Router 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:12:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The open-source Claude Code Model Router provides a local API gateway so developers can run Claude Code sessions through third-party model providers, including the Seed provider with the 2.1-pro variant configured as th…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The project ships a runnable npm/CLI gateway and repository artifact that exposes commands such as init/start/claude and lists Seed aliases so users can select Seed 2.1 Pro for Claude Code-style coding workflows.

模型作用:Seed-2.1-Pro-Preview 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Seed 2.1 Pro is configured as the Doubao Seed provider's default 2.1-pro variant with Volcengine Ark-compatible endpoint settings, 262144 max tokens, and 262144 context window, enabling the gateway to route coding-agent…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from an open-source integration repository rather than a formal customer story. The repository names the production/API variant as Doubao Seed 2.1 Pro with model_id doubao-seed-2-1-pro-260628 and aliases see…

原始记录:Claude Code Model Router integrates Doubao Seed 2.1 Pro for third-party Claude Code coding sessions

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Aero 使用 Qwen3.6 Max Preview 处理可玩交互原型构建

Aero · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:youtube 最后复核:2026-06-27T10:20:16Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Aero 公开的游戏与交互原型案例,来源为 视频证据,复核于 2026-06-27T10:20:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕可玩交互原型构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Aero tested Qwen3.6 Max Preview in Qwen Studio with concrete generation prompts: an interactive transformer-teaching webpage, an interactive Pokémon Pokédex webpage, and a playable Flappy Bird-style browser game.

公开产物:公开材料提供视频记录、页面说明或可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The video shows generated web artifacts: a transformer explainer with animations for parallel processing, positional encodings and attention weights; a Flappy Bird game with polished graphics and animation but difficult…

模型作用:Qwen3.6 Max Preview 在该案例中承担可玩交互原型构建相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3.6 Max Preview generated the HTML/web app implementations and UI content from prompts, including educational animations, game visuals/logic, formatted code sections, and dashboard-style application structure.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public creator demo rather than a production customer deployment; evidence is accepted because the video transcript explicitly names Qwen3.6 Max Preview, Qwen Studio, the user/channel, concrete tasks, and visible genera…

原始记录:Aero used Qwen3.6 Max Preview in Qwen Studio to generate interactive web prototypes

已有真实案例 游戏与交互原型视频证据A 类可核验real_case auto_approved 进入模型卡精选

ModelScope / Alibaba Cloud 使用 Qwen3 Max Thinking (Preview) 处理软件工程任务执行

ModelScope / Alibaba Cloud · Qwen3 Max Thinking (Preview)

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking (Preview) 来源平台:official_web 最后复核:2026-06-27T10:20:12Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

ModelScope / Alibaba Cloud 公开的代码代理与软件工程案例,来源为 官方页面,复核于 2026-06-27T10:20:12Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:ModelScope publishes a public model page at the exact Qwen/Qwen3-Max-Thinking path, making the Qwen3 Max Thinking model artifact discoverable through ModelScope's model exploration, inference, training, deployment and a…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public ModelScope artifact URL is reachable and exposes the exact Qwen/Qwen3-Max-Thinking model route as a hosted model repository entry for users to find and use through the ModelScope platform.

模型作用:Qwen3 Max Thinking (Preview) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3-Max-Thinking is the reasoning model being distributed through the ModelScope artifact page, contributing Alibaba Qwen's long-context thinking and advanced reasoning capabilities to ModelScope users' exploration an…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Provider/platform hosting case rather than a downstream customer story; accepted because the public artifact URL is reachable, exact model identity is bound in the URL, and ModelScope is an Alibaba-operated model hostin…

原始记录:ModelScope hosts the Qwen/Qwen3-Max-Thinking public model artifact for exploration and deployment workflows

已有真实案例 代码代理与软件工程官方页面A 类可核验real_case auto_approved 进入模型卡精选

tmousemy881 / benjiyaya/HeartMuLa_C… 使用 Qwen3 Max Thinking 处理真实任务执行

tmousemy881 / benjiyaya/HeartMuLa_ComfyUI · Qwen3 Max Thinking

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking 来源平台:github 最后复核:2026-06-27T10:12:09Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

tmousemy881 / benjiyaya/HeartMuLa_ComfyUI 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T10:12:09Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:HeartMuLa_ComfyUI 的用户 tmousemy881 在公开 issue #74 中提交 HeartMuLa Music Style Selector 节点方案,要求替换 custom_nodes/HeartMuLa_ComfyUI 目录下的 __init__.py,并放置 style_mapping.txt,将中文风格/乐器提示词映射为 HeartMuLa 标准标签且支持多选;issue 明确说明该工作是在 Qwen3-Max-Thinking 辅助下开发完成。

公开产物:公开产物包含可下载的 __init__.py 和 style_mapping.txt 附件,以及一张节点界面截图;style_mapping.txt 已核验可访问并包含中文到英文标签映射,__init__.py 附件返回 text/x-python,说明该案例产生了可安装到 ComfyUI 自定义节点目录的代码与配置文件。

模型作用:Qwen3 Max Thinking 的贡献是辅助无代码背景用户完成 ComfyUI 自定义节点代码与风格映射文件的设计/实现,把音乐生成提示词选择流程产品化为 HeartMuLa Music Style Selector 节点,而不是仅用于 benchmark 或教程演示。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub issue 的 JSON-LD/页面正文,明确写有 developed with the assistance of Qwen3-Max-Thinking;产物是 issue 附件而非已合并 PR,仓库主分支 README 未显示该节点,因此按真实用户提交的公开可访问代码产物收录。

原始记录:HeartMuLa_ComfyUI 用户用 Qwen3 Max Thinking 辅助开发音乐风格选择器节点

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Bijan Bowen 使用 Qwen3.6 Max Preview 处理浏览器 3D 世界构建

Bijan Bowen · Qwen3.6 Max Preview

A
厂商:Qwen / Alibaba 模型:Qwen3.6 Max Preview 来源平台:youtube 最后复核:2026-06-27T10:20:16Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 97/100

A 高可信 · 社区公开记录

原始证据1 个公开产物复核通过社区公开记录

Bijan Bowen 公开的3D 与 Web 交互案例,来源为 视频证据,复核于 2026-06-27T10:20:16Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕浏览器 3D 世界构建的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Bijan Bowen selected Qwen3.6 Max Preview in the Qwen web chat interface and also connected it through OpenCode, then asked it to create a browser-based operating system with five applications, two functional 3D games in…

公开产物:公开材料提供视频记录、页面说明或可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The video reviews the generated browser OS artifact, including login flow, notifications, date/time, web browser, settings, media player, file manager, terminal commands, two game apps, and a novel 3D desktop-mode speci…

模型作用:Qwen3.6 Max Preview 在该案例中承担浏览器 3D 世界构建相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3.6 Max Preview produced multi-file/code-heavy prototype outputs from long prompts, planned and generated browser OS functionality, UI elements, games, and OpenCode-driven implementation artifacts for the creator to…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Public hands-on evaluation rather than production deployment; evidence is accepted because the transcript explicitly identifies Qwen3.6 Max Preview, the user/channel, the coding environment, concrete artifact requiremen…

原始记录:Bijan Bowen used Qwen3.6 Max Preview to build a browser OS and game prototypes

已有真实案例 3D 与 Web 交互视频证据A 类可核验real_case auto_approved 进入模型卡精选

ByteDance Seed 使用 Seed2.1 Turbo 处理智能体流程编排

ByteDance Seed · Seed2.1 Turbo

A
厂商:ByteDance Seed 模型:Seed2.1 Turbo 来源平台:official_web 最后复核:2026-06-27T10:19:51Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

ByteDance Seed 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T10:19:51Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The official Seed2.1 page lists an Agent Practice across Tools & Environments showcase for the released Seed2.1 family, including Seed2.1 Turbo. The task is continuous office execution across multiple tools, environment…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The official page exposes a separate public HTML artifact for the showcase, demonstrating a cross-tool agent experience that can continue office work across interfaces and environments.

模型作用:Seed2.1 Turbo 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:Seed2.1 Turbo contributes computer-use-style task planning, cross-environment state tracking, GUI/non-GUI tool switching, and multi-step execution reliability, matching the page's description of Seed2.1 as an agent for …

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:First-party official showcase and public artifact, not an independent external deployment. The binding to Turbo is via the Seed2.1 model page that states the family has Pro and Turbo variants and publishes Turbo benchma…

原始记录:ByteDance Seed showcases Seed2.1 Turbo for cross-tool office agent execution

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

easyicetrusty / BerriAI LiteLLM + L… 使用 Qwen3 Max Thinking 处理真实任务执行

easyicetrusty / BerriAI LiteLLM + LibreChat deployment · Qwen3 Max Thinking

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking 来源平台:github 最后复核:2026-06-27T10:23:35Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

easyicetrusty / BerriAI LiteLLM + LibreChat deployment 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T10:23:35Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:用户 easyicetrusty 在公开 LiteLLM issue 中说明,其私有 dorx LibreChat 部署通过 LiteLLM proxy、DashScope/Alibaba Cloud OpenAI-compatible 接口调用 `qwen/qwen3-max-thinking` 等 Qwen thinking 模型;任务是在流式 chat-completion 中处理 provider 内容审核 400,并通过 content_policy_fallbacks 路由到 tier-2 模型,而不是让会话 500。

公开产物:issue 给出可复现路径、环境和补丁:`litellm[proxy]` main-latest、LibreChat streaming consumer、DashScope provider、模型 `qwen/qwen3-max-thinking`;作者报告当前生产环境中已将补丁以 bind-mounted vendored file 方式运行,Streaming DashScope content-policy 400s on Qwen thinking models can now fall through to a tier-2 model instead of returning a misleading `Fallbacks=None` hard error。

模型作用:Qwen3 Max Thinking 在该案例中是实际部署的 DashScope/Qwen thinking 后端模型;它的内容策略 400 与流式返回行为触发了 LiteLLM fallback 代码路径缺陷,使用户能够围绕真实 Qwen3 Max Thinking 流式会话定位、验证并提出对 ContentPolicyViolationError 的 MidStreamFallbackError 处理修复。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub issue 的结构化正文,明确列出 Environment 中的 model=`qwen/qwen3-max-thinking`、DashScope provider、LibreChat streaming consumer 和生产 workaround;未公开私有 dorx LibreChat 实例或具体被拦截 prompt,因此 artifact 使用同一公开 issue。不是 benchmark、教程或发布文。

原始记录:LibreChat/LiteLLM 用户用 Qwen3 Max Thinking 复现并修补 DashScope 流式内容策略回退问题

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Minh Ha-Duong / CIRED-CNRS AEDIST p… 使用 Qwen3 Max Thinking (Preview) 处理软件工程任务执行

Minh Ha-Duong / CIRED-CNRS AEDIST project · Qwen3 Max Thinking (Preview)

A
厂商:Qwen / Alibaba 模型:Qwen3 Max Thinking (Preview) 来源平台:github 最后复核:2026-06-27T10:20:21Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Minh Ha-Duong / CIRED-CNRS AEDIST project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:20:21Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The AEDIST project, a technical feasibility report for AI-driven energy data integration, used Qwen3 Max Thinking as one of the models to audit a draft technical report and identify the strongest inconsistency, weakest …

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public audit artifact records the model run as qwen/qwen3-max-thinking on 2026-05-01, with token counts and wall time, and includes a concrete four-part critique: it flags an inconsistency between the claimed satura…

模型作用:Qwen3 Max Thinking (Preview) 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen3 Max Thinking contributed analytical review of the report by reading the paper context and producing specific critique points for methodological consistency, empirical support, and claim selection, which were saved…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public GitHub repository and raw artifact is reachable. The artifact records the model slug as qwen/qwen3-max-thinking; this candidate maps it to the Atlas model Qwen3 Max Thinking (Preview) because the ta…

原始记录:AEDIST technical report used Qwen3 Max Thinking to audit claims in an energy-data AI benchmark paper

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

LobeHub / LobeChat 使用 Claude 2.0 处理软件工程任务执行

LobeHub / LobeChat · Claude 2.0

A
厂商:Anthropic / Claude 模型:Claude 2.0 来源平台:github 最后复核:2026-06-27T10:24:52Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

LobeHub / LobeChat 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:24:52Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:LobeHub's open-source LobeChat application maintains a model-bank configuration that includes Anthropic Claude 2.0 (id claude-2.0, generation claude-2) as a chat model option, enabling users of the product to run genera…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public LobeChat repository is an accessible AI chat application artifact. Its model-bank source defines a Claude 2.0 chat entry with 100,000 context-window tokens, 4,096 max output tokens, July 11 2023 release metad…

模型作用:Claude 2.0 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Claude 2.0 contributes the long-context text understanding and generation capability used by LobeChat when users select that model for assistant-style chat, document/context-heavy prompting, and other general conversati…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from the live open-source product repository rather than a customer story. The product is multi-model and the specific file is a model catalog/configuration, so it proves Claude 2.0 was exposed as a supporte…

原始记录:LobeHub LobeChat exposed Claude 2.0 as a selectable model for open-source AI chat workflows

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

elmorshedy-del / Virona-ShawQ-Dashb… 使用 GPT-5.1 (high) 处理研究分析和报告生成

elmorshedy-del / Virona-ShawQ-Dashboard · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:github 最后复核:2026-06-27T10:25:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

elmorshedy-del / Virona-ShawQ-Dashboard 公开的研究与报告生成案例,来源为 公开代码库,复核于 2026-06-27T10:25:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕研究分析和报告生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The dashboard's Creative Intelligence chat added an OpenAI GPT-5.1 model option and a reasoning_effort setting defaulting to high for deeper reasoning and creative strategy sessions.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The pull request changes the React settings UI, database migration, route handling, and OpenAI service so Creative Intelligence chat requests can persist and send GPT-5.1 with high reasoning effort.

模型作用:GPT-5.1 (high) 在该案例中承担研究分析和报告生成相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 high is used as the deeper-reasoning model behind creative analysis conversations, intended to produce stronger strategic reasoning and creative connections than the existing model choices.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public product repository pull request and patch; it demonstrates integration and intended use, but does not include production traffic metrics or end-user output samples.

原始记录:Virona ShawQ Dashboard adds GPT-5.1 high-effort reasoning to Creative Intelligence chat

已有真实案例 研究与报告生成公开代码库A 类可核验real_case auto_approved 进入模型卡精选

geminii01 / UpThink project team 使用 Solar Pro 2 处理软件工程任务执行

geminii01 / UpThink project team · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:GitHub 最后复核:2026-06-27T10:25:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

geminii01 / UpThink project team 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:25:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Build an Obsidian-focused personal knowledge management service that processes markdown notes, generates image alt text, recommends tags, connects related notes, and helps split unstructured documents.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public repository for UpThink with README describing the service, architecture, usage flow, and Solar Pro 2-powered features for automating repetitive knowledge-organization work.

模型作用:Solar Pro 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Solar Pro 2 is explicitly named as the LLM used for language-understanding tasks: generating approximately 50-word image alt text from parsed image content, recommending tags from note content, and supporting note organ…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub README for an Upstage AI Ambassador project; public repo is accessible, but no independent production deployment metrics were found.

原始记录:UpThink uses Solar Pro 2 to automate Obsidian personal knowledge management

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

TheAnswerManIsHere / Overhypeme 使用 GPT-5.1 (high) 处理多模态内容处理

TheAnswerManIsHere / Overhypeme · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:github 最后复核:2026-06-27T10:25:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

TheAnswerManIsHere / Overhypeme 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-27T10:25:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Overhypeme added GPT-5.1, GPT-5.2, and GPT-5.4 mini to the model dropdown for AI Image Style Prompt and AI Video Motion Prompt generation, with a per-prompt reasoning effort lever including high.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The pull request routes both scene-prompt and video-motion generators through shared OpenAI chat parameter logic that sends reasoning_effort and max_completion_tokens for GPT-5 reasoning models while omitting incompatib…

模型作用:GPT-5.1 (high) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 high can be selected by admins to generate more capable and subtle meme/image/video prompt text, using higher reasoning effort when quality is preferred over cost or latency.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public repository PR and patch; the high setting is exposed as an operator-selectable option rather than forced for every request.

原始记录:Overhypeme enables GPT-5.1 high-effort reasoning for AI image style and video motion prompts

已有真实案例 多模态生成与理解公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Bondartsov 使用 Grok 4.20 0309 处理代码审查和测试生成

Bondartsov · Grok 4.20 0309

A
厂商:xAI / Grok 模型:Grok 4.20 0309 来源平台:GitHub 最后复核:2026-06-27T10:26:24Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Bondartsov 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T10:26:24Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:grok-critic-mcp 是公开 MCP 服务器,README 明确说明它通过 Polza.AI / Responses API 调用 xAI 的 grok-4.20-multi-agent,为 AI 编程代理提供 critic_review、architecture_review、security_audit、critic_followup、health/config/restart 等工具,并支持 4/16 agents 的 reasoning effort 映射。

公开产物:仓库提供可安装 Python 包、FastMCP server、配置、技能文档和测试;用户可把它接入 Kilo Code、Cursor、Claude Code 等客户端,对代码片段或文件执行深度代码审查、架构评审和安全审计,并返回带模型、agents、effort 元数据的审查结果。

模型作用:Grok 4.20 Multi-Agent 作为 MCP 服务背后的审查模型,使用多 reasoning agents 并行分析代码、架构和安全风险,再形成共识式批评与改进建议,是该工具的核心判断与生成能力来源。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据明确写 grok-4.20-multi-agent / x-ai/grok-4.20-multi-agent;未在仓库标题中写 0309 后缀,但可绑定到 Grok 4.20 0309 multi-agent 家族。公开仓库可访问,属于具体工具产物而非教程、benchmark 或综述。

原始记录:Bondartsov 的 grok-critic-mcp 用 Grok 4.20 Multi-Agent 做代码审查、架构分析和安全审计

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选

bnyoun / BizTone Converter project 使用 Solar Pro 2 处理软件工程任务执行

bnyoun / BizTone Converter project · Solar Pro 2

A
厂商:Upstage / Solar 模型:Solar Pro 2 来源平台:GitHub 最后复核:2026-06-27T10:25:34Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

bnyoun / BizTone Converter project 公开的代码代理与软件工程案例,来源为 公开代码库,复核于 2026-06-27T10:25:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕软件工程任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Provide a web service that transforms rough Korean drafts or key points into polished business language tailored to recipients such as managers, colleagues, customers, or team members.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Public FastAPI plus vanilla JavaScript repository for a business-tone converter, with README documenting text input, recipient persona selection, tone conversion, copy-to-clipboard output, and Vercel-oriented deployment.

模型作用:Solar Pro 2 在该案例中承担软件工程任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Upstage Solar-Pro2 is explicitly listed as the LLM, via LangChain, used for precise tone transformation into professional business communication across four recipient personas.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is from a public GitHub README and source artifact; the deployment link was not separately verified, so the repository is used as the public artifact.

原始记录:BizTone Converter uses Solar-Pro2 to rewrite rough drafts into business communication

已有真实案例 代码代理与软件工程公开代码库A 类可核验real_case auto_approved 进入模型卡精选

latitude-dev / latitude-llm 使用 GPT-5.1 (high) 处理知识检索和问答

latitude-dev / latitude-llm · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:github 最后复核:2026-06-27T10:25:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

latitude-dev / latitude-llm 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T10:25:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕知识检索和问答的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Latitude's LLM platform fixed Azure provider option routing so prompt configurations using model gpt-5.1 with reasoning_effort set to high are delivered under the OpenAI provider-options key expected by the Vercel AI SD…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The patch adds regression coverage showing a Latitude config with model gpt-5.1 and reasoning_effort high is transformed into providerOptions.openai.reasoningEffort = high, preserving Azure resource metadata.

模型作用:GPT-5.1 (high) 在该案例中承担知识检索和问答相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 high supplies the reasoning-heavy execution mode for Latitude prompt runs on Azure OpenAI; the change ensures the model actually receives the high-effort instruction instead of silently dropping it.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a targeted bug-fix PR with explicit gpt-5.1 high configuration in tests; it verifies request routing rather than publishing user-facing generated outputs.

原始记录:Latitude fixes Azure GPT-5.1 high-effort prompt routing for its LLM platform

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Microsoft / Visual Studio Code 使用 GPT-5.1 (high) 处理智能体流程编排

Microsoft / Visual Studio Code · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:github 最后复核:2026-06-27T10:25:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Microsoft / Visual Studio Code 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T10:25:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕智能体流程编排的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The VS Code Copilot extension added configurable Azure BYOK Entra authentication and model capability metadata so Azure OpenAI deployments such as gpt-5.1 can expose a Thinking Effort picker and send selected high effor…

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The PR adds configuration schema, Azure provider capability resolution, tests for gpt-5.1 with supported reasoning efforts none/low/medium/high, and request handling that preserves the configured high effort.

模型作用:GPT-5.1 (high) 在该案例中承担智能体流程编排相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 high acts as the reasoning mode behind developer-assistant chat/code workflows in VS Code when a user configures an Azure BYOK GPT-5.1 deployment and selects high thinking effort.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a Microsoft public PR and patch. It is a product-integration case for a configurable model deployment, not a public customer story with usage volume metrics.

原始记录:VS Code forwards Azure BYOK GPT-5.1 high reasoning effort through the Responses API

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

iii-hq / workers 使用 GPT-5.1 (high) 处理真实任务执行

iii-hq / workers · GPT-5.1 (high)

A
厂商:OpenAI 模型:GPT-5.1 (high) 来源平台:github 最后复核:2026-06-27T10:25:47Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

iii-hq / workers 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T10:25:47Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:iii-hq implemented a standalone OpenAI provider worker with per-family reasoning-effort mapping, including GPT-5.1 support for none, low, medium, and high effort, and request-shape handling for chat usage.

公开产物:公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The PR introduces a Rust provider-openai worker, reasoning-effort mapping logic, manifest/config tests, and validation notes including a high-thinking GPT-5.1 call that answered a bat-and-ball reasoning check correctly …

模型作用:GPT-5.1 (high) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:GPT-5.1 high is used as the high-reasoning OpenAI path in the worker, enabling downstream chat requests to request stronger reasoning behavior from GPT-5.1 instead of treating all models with a single effort vocabulary.

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Evidence is a public implementation PR with explicit GPT-5.1 high handling and validation notes; it is primarily infrastructure usage rather than a polished public demo page.

原始记录:iii-hq Workers adds an OpenAI provider worker with GPT-5.1 high-effort reasoning support

已有真实案例 真实任务执行公开代码库A 类可核验real_case auto_approved 进入模型卡精选

shubhamsinghrathore 使用 Claude Mythos 5 处理知识检索和问答

shubhamsinghrathore · Claude Mythos 5

A
厂商:Anthropic / Claude 模型:Claude Mythos 5 来源平台:github 最后复核:2026-06-27T10:32:15Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

shubhamsinghrathore 公开的知识库与检索问答案例,来源为 公开代码库,复核于 2026-06-27T10:32:15Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:创建一组面向 Claude 的 DevOps/平台工程调试技能,用于 CI/CD 失败分析、Kubernetes 故障诊断与安全审计、Infrastructure-as-Code/Terraform/Terragrunt/Crossplane 架构与故障排查等场景。仓库 README 明确写明作者 created Claude skills with the use of Mythos 5 model,仓库内公开包含 cicd-failure-analyzer、k8s-guardian、iac-architect 等完整技能文档。

公开产物:公开产物为 GitHub 仓库 shubhamsinghrathore/CLAUDE-SKILLS;仓库包含可直接阅读和复用的 Claude Skills markdown 文件:cicd-failure-analyzer.md、k8s-guardian.md、iac-architect.md,每个文件都有技能名、description、reasoning_effort、compatibility 和分阶段诊断/交付流程。

模型作用:README 将 Mythos 5 model 明确列为创建这些 Claude skills 的使用模型;模型贡献体现在把 CI/CD、Kubernetes、IaC 等工程排障知识组织为结构化、可触发的 Claude Skills,包括诊断树、证据要求、风险评分、回滚计划和输出格式。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据来自公开 GitHub 仓库 README 与仓库文件,已用 git clone/raw README 核验可访问;模型归因是作者自述,README 写作较简短且未提供独立第三方复核。该条不是 benchmark、教程、集合页或发布综述,artifact 是实际技能文档仓库。

原始记录:shubhamsinghrathore 用 Claude Mythos 5 创建 DevOps 调试类 Claude Skills

已有真实案例 知识库与检索问答公开代码库A 类可核验real_case auto_approved 进入模型卡精选

ByteDance Seed 使用 Seed2.1 Turbo 处理多模态内容处理

ByteDance Seed · Seed2.1 Turbo

A
厂商:ByteDance Seed 模型:Seed2.1 Turbo 来源平台:official_web 最后复核:2026-06-27T10:29:26Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

ByteDance Seed 公开的多模态生成与理解案例,来源为 官方页面,复核于 2026-06-27T10:29:26Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The official Seed2.1 page lists a concrete showcase named 'Long Movies to Narrated Shorts' under multimodal agent applications. The task is to understand a long-form video, select and edit highlights, and generate a nar…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public page exposes a reachable MP4 artifact for the narrated-shorts workflow, demonstrating a long-movie-to-short-video result rather than a text-only answer or benchmark score.

模型作用:Seed2.1 Turbo 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Seed2.1 Turbo contributes multimodal video understanding, temporal reasoning over long clips, content selection, narrative planning, and generation/editing support so the workflow can transform long video input into a u…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:First-party official showcase rather than an independent customer story. The case binds to Turbo through the Seed2.1 family page, which explicitly states the family includes Pro and Turbo, and uses the separate official…

原始记录:ByteDance Seed showcases Seed2.1 Turbo for turning long movies into narrated short videos

已有真实案例 多模态生成与理解官方页面A 类可核验real_case auto_approved 进入模型卡精选

ByteDance Seed 使用 Seed2.1 Turbo 处理多模态内容处理

ByteDance Seed · Seed2.1 Turbo

A
厂商:ByteDance Seed 模型:Seed2.1 Turbo 来源平台:official_web 最后复核:2026-06-27T10:31:20Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

ByteDance Seed 公开的多模态生成与理解案例,来源为 官方页面,复核于 2026-06-27T10:31:20Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The official Seed2.1 page's multimodal showcase lists a case titled 'Long Movies to Narrated Shorts' for the released Seed2.1 family, whose model table includes Seed2.1 Turbo. The task is to understand, edit, and narrat…

公开产物:公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The official page exposes a public MP4 artifact for the showcase, demonstrating a long-video workflow that turns source video content into a narrated short/highlight output rather than a text-only response.

模型作用:Seed2.1 Turbo 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:Seed2.1 Turbo contributes long-video understanding, multimodal perception, temporal reasoning, content selection, editing-planning, and narration-generation capabilities described for the Seed2.1 family, enabling end-to…

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:First-party official showcase and reachable video artifact; it binds to Seed2.1 Turbo through the same official page's released Pro/Turbo model table rather than an independent external customer deployment.

原始记录:ByteDance Seed showcases Seed2.1 Turbo for turning long movies into narrated shorts

已有真实案例 多模态生成与理解官方页面A 类可核验real_case auto_approved 进入模型卡精选

火山方舟(Volcengine Ark) 使用 Seed2.1 Pro 处理智能体流程编排

火山方舟(Volcengine Ark) · Seed2.1 Pro

A
厂商:ByteDance Seed 模型:Seed2.1 Pro 来源平台:official_web 最后复核:2026-06-27T10:34:04Z 证据快照:待快照 · 0 / 2 个证据目标 归档说明:已列入归档队列,等待抓取 2 个证据目标。
证据可信度 99/100

A+ 完整链路 · 官方/客户故事

原始证据1 个公开产物复核通过官方/客户故事

火山方舟(Volcengine Ark) 公开的智能体工作流案例,来源为 官方页面,复核于 2026-06-27T10:34:04Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Seed2.1 官方项目页在体验入口中提供火山方舟公开体验链接和 API 入口,并将体验链接参数明确绑定到 modelId=doubao-seed-2-1-pro-260628;面向开发者和业务用户的任务是在火山方舟中直接试用或通过 API 调用 Seed2.1 Pro,完成通用 Agent、多步骤办公、代码工程、视觉/长上下文理解等生产力任务。

公开产物:火山方舟体验中心链接可公开访问并加载,URL 中直接携带 doubao-seed-2-1-pro-260628 模型标识;同一项目页还提供模型详情/API 入口 Id=doubao-seed-2-1-pro,形成可核验的公开产物与调用入口。

模型作用:Doubao-Seed-2.1-Pro 作为火山方舟体验中心和 API 入口绑定的具体模型,向开发者提供 Seed2.1 Pro 的 Agent 规划、跨工具任务推进、代码工程交付、多模态理解与长上下文处理能力,而不是仅停留在发布说明或 benchmark 展示。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:一手官方证据和一方平台产物;这是平台级真实上线/可调用案例,不是第三方客户故事。已核验 artifact_url HTTP 200,API 控制台详情页需登录但公开体验中心可访问。

原始记录:火山方舟上线 Doubao-Seed-2.1-Pro 体验中心与 API,用于开发者真实调用 Agent/Coding 模型能力

已有真实案例 智能体工作流官方页面A 类可核验real_case auto_approved 进入模型卡精选

AlphaOne LLC 使用 Grok 4.20 0309 v2 处理智能体流程编排

AlphaOne LLC · Grok 4.20 0309 v2

A
厂商:xAI / Grok 模型:Grok 4.20 0309 v2 来源平台:GitHub 最后复核:2026-06-27T10:42:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

AlphaOne LLC 公开的智能体工作流案例,来源为 公开代码库,复核于 2026-06-27T10:42:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:ai-memory-a2a-v0.7.0 的 scripts/grok_driver.py 是围绕 xAI Chat Completions 的标准库驱动,文件说明 S67 和 S68 场景会用真实的 grok-4.20-0309-reasoning 模型执行推理,并默认从 XAI_MODEL 读取该 exact model id。

公开产物:公开仓库提供 grok_chat 调用封装、reasoning/token usage 提取和命令行 smoke test;驱动返回 {text, reasoning, model, usage} 结构,供 A2A memory 场景把提示交给 Grok 4.20 0309 Reasoning 并记录模型输出与用量。

模型作用:Grok 4.20 0309 Reasoning 被作为 S67/S68 真实推理后端,负责根据场景 prompt 生成回答,并通过 usage/reasoning token 信息支持后续记忆与评估流程。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据文件直接写明默认模型 grok-4.20-0309-reasoning,并说明具体 S67/S68 场景使用它驱动实际 reasoning model;这是公开工程仓库证据,未独立验证外部生产流量。

原始记录:AlphaOne 的 ai-memory-a2a 用 Grok 4.20 0309 Reasoning 驱动记忆场景推理

已有真实案例 智能体工作流公开代码库A 类可核验real_case auto_approved 进入模型卡精选

Lumiwealth / Lumibot 使用 Grok 4.20 0309 v2 处理代码审查和测试生成

Lumiwealth / Lumibot · Grok 4.20 0309 v2

A
厂商:xAI / Grok 模型:Grok 4.20 0309 v2 来源平台:GitHub 最后复核:2026-06-27T10:42:00Z 证据快照:已快照 · 2 / 2 个证据目标 归档说明:已完成快照:2 / 2 个证据目标,可用本地 manifest 复核证据指纹。
证据可信度 100/100

A+ 完整链路 · 代码仓库证据

原始证据1 个公开产物复核通过代码仓库证据

Lumiwealth / Lumibot 公开的代码审查与测试案例,来源为 公开代码库,复核于 2026-06-27T10:42:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。

任务:Lumibot 的 agent_m2_liquidity_grok.py 是 M2 Liquidity Strategy 的 xAI Grok 版本,文件说明 default_model 使用 xai/grok-4.20-0309-reasoning,并通过 agent 工具抓取 FRED 宏观数据,判断 M2 流动性扩张或收缩后在 TQQQ 与 SHV 之间切换。

公开产物:公开示例代码创建名为 m2_analyst 的 agent,使用 FRED/ALFRED 数据工具和 YahooDataBacktesting 执行 2024-01-01 到 2025-01-01 的回测;策略始终持有 TQQQ 或 SHV,并把 Grok 的宏观分析转成交易动作。

模型作用:Grok 4.20 0309 Reasoning 作为策略 agent 的默认模型,负责读取宏观数据、比较近月与 3-6 个月前的 M2 趋势,生成买入 TQQQ 或 SHV 的决策依据。

A 类理由:有具体使用者、具体任务、公开原始证据和可访问产物,可公开核验。

风险边界:证据文件直接绑定 xai/grok-4.20-0309-reasoning;这是开源策略示例和回测产物,不应包装成真实资金收益或投资建议。

原始记录:Lumiwealth 的 Lumibot 示例用 Grok 4.20 0309 Reasoning 做 M2 流动性交易策略分析

已有真实案例 代码审查与测试公开代码库A 类可核验real_case auto_approved 进入模型卡精选