跳到正文
@kimmonismus· @kimmonismus · X·· 2026-06-03AI 评分56
AI 导读

微软 MAI 技术报告被解读为一款 1T 参数、35B 激活、在 33.5T token 上训练的模型,训练中未使用合成数据,也未从此前模型蒸馏。解读指出,该模型的推理、智能体行为与工具使用能力全部在后训练阶段习得,没有冷启动,这种做法难度更高、需要更多迭代才能达到 SOTA。报告还给出了各轮迭代的精确 MFU 以及完整的缩放配方(scaling ladder recipe)。

正文

Fantastic in depth guide about Microsoft MAI by @eliebakouch

tl;dr about the model: Respect where respect is due.

-zero synthetic data or distillation from previous models.

-1T model with 35B active, trained on 33.5T tokens

引用elie (@eliebakouch)@eliebakouch
microsoft MAI tech report is a gold mine, one of the most transparent for a model at this scale. this model uses zero synthetic data or distillation from previous models. this means reasoning, agentic behavior, tool use are all learned fully during post-training with no cold start. bold choice that makes it harder and requires more iterations to reach sota, but you get FULL control over your model series and it proves they are serious about being a frontier lab. the tech report is insanely detailed and precise about numbers. to give an example, they give the exact MFU across all the iterations of the model, with the exact changes etc. they also share the full scaling ladder recipe, to my knowledge this is the first time i've seen this in a tech report at this scale let's look at all of this in this likely very long thread 🧵
在 X 查看被引用的帖子

来源:@kimmonismus · x.com