跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 2026-09-03AI 评分37
AI 导读

NVIDIA Research 🚀 产出过一些出色的研究,比如 LatentMoE(用于 Kimi K3)和 GatedDeltaNets(用于 Qwen)。但在端到端前沿训练方面,NVIDIA 的官僚文化产出了像 Nemotron3 Ultra 这样令人尴尬的模型。 尽管 NVIDIA Research 拥有出色的人才,Nemotron3 Ultra 拥有 550B 总参数(55B 激活),却被所有中国模型碾压,甚至包括 Qwen3.8 27B 参数——后者的参数量少了约 20 倍。

正文

NVIDIA Research 🚀 has produced some great research, like LatentMoE (used in Kimi K3) and GatedDeltaNets (used in Qwen). But for e2e frontier training, NVIDIA's bureaucratic culture has produced embarrassing models like Nemotron3 Ultra.

Despite NVIDIA Research having amazing talent, Nemotron3 Ultra, with 550B total params (55B active), is getting mogged by all the Chinese models, including even Qwen3.8 27B parameters, which has ~20x fewer parameters.

来源:@SemiAnalysis_ · x.com