跳到正文
原文
Google Developers Blog·· 8 天前AI 评分50

Google 用 MaxText 在 TPU 上复现 Ai2 的 Olmo 3 7B 预训练与中期训练

Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

AI 导读

Google 团队在 MaxText/JAX 和 Google Cloud TPU 上从零复现了 Ai2 的 Olmo 3 7B,覆盖 stage-1 预训练(约5.93T token)和 stage-2 中期训练 anneal,held-out C4 loss 与 8 项下游任务均与 Ai2 参考持平。

来源:Google Developers Blog · developers.googleblog.com