Chinese Brief
中文案例导读
hiyouga/LLaMA-Factory GitHub user community 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
CASE EVIDENCE / A RECORD
hiyouga/LLaMA-Factory GitHub user community 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。
原始记录:LLaMA-Factory user attempts SFT fine-tuning of Qwen1.5-110B-Chat with DeepSpeed ZeRO-3 on 4×A40
Chinese Brief
hiyouga/LLaMA-Factory GitHub user community 公开的真实任务执行案例,来源为 公开代码库,复核于 2026-06-27T07:46:08Z。 Model Atlas 将它标记为 A 类证据,因为它同时具备具体使用者、具体任务、公开原始证据和可访问产物。 Model Atlas 不把 benchmark、教程、发布说明或集合页包装成真实案例。
任务
这是一个围绕真实任务执行的真实任务,公开材料可以回溯到具体使用者和具体产物。 原始资料写作:A user configured LLaMA-Factory SFT with DeepSpeed ZeRO-3 offload to fine-tune ../../model/qwen/Qwen1.5-110B-Chat on a custom dataset named Kee_Instruction_NewEstabalish, using LoRA rank 128, cutoff_len 6000, and 4 A40 …
公开材料提供公开代码、README 或项目配置,可用于核验任务结果、项目形态和模型绑定关系。 原始资料写作:The public issue includes the training command, dataset/configuration, expected goal of using 4×A40 plus DeepSpeed ZeRO-3 to fine-tune Qwen1.5-110B, and the encountered device placement runtime error.
Qwen1.5 Chat 110B 在该案例中承担真实任务执行相关的生成、分析、编排或实现角色。 原始资料写作:Qwen1.5-110B-Chat was the base instruction/chat model being adapted through SFT/LoRA for the user's domain dataset.
当前判断基于公开材料;若产物下线、仓库变更或模型参与比例仅来自作者自述,需要在引用前重新复核。 原始资料写作:This is a fine-tuning troubleshooting artifact, not a successful public demo; however it documents a concrete organization/user workflow with exact model path, task, dataset, and training setup.