预训练并行策略与训练失败原因笔记
Notes on pretraining parallelisms and failed training runs.
阅读原文
本站未展示全文,请前往来源网站阅读。
AI 导读
预训练失败的两大主因是破坏因果性和引入偏差:专家路由中的 expert choice 与 token dropping 会让训练看到部署时不存在的信息,据传这解释了 Llama 4 表现不佳,Gemini 2 Pro 也曾受 token dropping 困扰。GPT-4 早期训练因在 all-reduce 等集合通信中使用 FP16,累加小梯度时数值精度丢失导致结果偏差达 10 倍。
来源:Dwarkesh Patel · dwarkesh.com