DATA CUT 2026-07-27 116 活跃模型 680 A 类案例

CASE EVIDENCE / A RECORD

Moonshot AI Research Team 使用 Kimi K2.5 处理代码审查和测试生成

Moonshot AI Research Team 公开的代码审查与测试案例,来源为 官方页面,复核于 2026-06-26T22:55:00Z。

原始记录:K2.5 Powering WorldVQA Visual Knowledge Benchmark Research

A

Chinese Brief

中文案例导读

Moonshot AI Research Team 公开的代码审查与测试案例,来源为 官方页面,复核于 2026-06-26T22:55:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。

厂商Kimi / Moonshot AI
模型Kimi K2.5
任务类型代码审查与测试
审核状态auto_approved

任务

真实任务背景

这是一个围绕代码审查和测试生成的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:Evaluate Kimi K2.5 on WorldVQA, a novel benchmark testing atomic visual knowledge of frontier models across 9 task categories (People, Objects, Culture, Geography, Transportation, Brands, Sports, Entertainment, Location…

代码审查与测试官方页面A 类可核验real_case
公开产物

公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Kimi K2.5 achieved 46.3% accuracy on WorldVQA Overall Accuracy, ranking second only to Gemini-3-pro (47.4%) and outperforming Claude-opus-4.5 (36.8%), GPT-5.2 (28.0%), Qwen3-VL-235B (23.5%), and all other evaluated mode…

模型作用

Kimi K2.5 在该案例中承担代码审查和测试生成相关的生成、分析、编排或实现角色。 原始资料写作:Kimi K2.5 demonstrates state-of-the-art visual knowledge understanding, achieving near-Gemini-3-pro performance on a challenging benchmark that tests real-world visual knowledge rather than OCR or simple recognition. Th…

风险边界

当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Published on Moonshot AI's official blog with arxiv paper reference. Benchmark data is from the model provider's own research but with open-source reproducible evaluation.