DATA CUT 2026-07-27 116 活跃模型 680 A 类案例

CASE EVIDENCE / A RECORD

LMSYS Org / SGLang Team 使用 DeepSeek R1 (Jan) 处理真实任务执行

LMSYS Org / SGLang Team 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T02:15:00Z。

原始记录:LMSYS/SGLang Team Deploys DeepSeek R1 on NVIDIA GB200 NVL72 with 3.8x/4.8x Throughput Gains

A

Chinese Brief

中文案例导读

LMSYS Org / SGLang Team 公开的真实任务执行案例,来源为 官方页面,复核于 2026-06-27T02:15:00Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。

厂商DeepSeek
模型DeepSeek R1 (Jan)
任务类型真实任务执行
审核状态auto_approved

任务

真实任务背景

这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:The SGLang team at LMSYS deployed and optimized DeepSeek R1 (and V3) inference on NVIDIA GB200 NVL72 hardware, implementing prefill-decode disaggregation with FP8 attention, NVFP4 MoE quantization, and large-scale exper…

真实任务执行官方页面A 类可核验real_case
公开产物

公开材料提供原始证据链接和可访问产物,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Achieved 26,156 input tokens/sec and 13,386 output tokens/sec per GPU on DeepSeek R1 for 2000-token input sequences — a 3.8x prefill and 4.8x decode throughput improvement over H100 baseline. With BF16 attention and FP8…

模型作用

DeepSeek R1 (Jan) 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:DeepSeek R1's Mixture-of-Experts architecture with fine-grained expert parallelism was the central target model; the 671B-parameter MoE design enabled the massive throughput gains on GB200 NVL72 through large-scale expe…

风险边界

当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:Blog post describes benchmark/optimization results from a research team rather than end-user production deployment; however LMSYS is a major systems research organization operating real GPU clusters, and the results are…