Chinese Brief
中文案例导读
LMSYS / SGLang Project 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
CASE EVIDENCE / A RECORD
LMSYS / SGLang Project 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。
原始记录:SGLang Day-0 Support for MiMo-V2-Flash with Optimized Inference
Chinese Brief
LMSYS / SGLang Project 公开的多模态生成与理解案例,来源为 公开代码库,复核于 2026-06-26T14:10:34Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
任务
这是一个围绕多模态内容处理的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:SGLang team (contributor acelyc111) integrated MiMo-V2-Flash as a first-class supported model in the SGLang LLM serving framework, implementing optimized Sliding Window Attention (SWA) execution, multi-layer MTP (Multi-…
公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:Merged PR #15207 with full day-0 support. Achieved 150 TPS per request decoding throughput even under 64K input tokens with batch size 16 per DP rank. Published detailed benchmarking results for prefill and decode phase…
MiMo-V2-Flash (Feb 2026) 在该案例中承担多模态内容处理相关的生成、分析、编排或实现角色。 原始资料写作:MiMo-V2-Flash's hybrid SWA+GA attention architecture and 3-layer MTP design enabled SGLang to demonstrate that inference-centric model design can achieve balanced throughput and latency. The model's architecture was spe…
当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:MiMo-V2-Flash was retired from Xiaomi MiMo Open Platform on 2026-06-18; model may no longer be available via official API. Historical use case.