NVIDIA 发布 Nemotron 3 Ultra,一款完全开源的 550B MoE 模型,激活参数 55B,权重、训练数据与完整配方全部公开。该模型采用混合 Mamba-Attention MoE 架构,NVIDIA 称其在长输出智能体任务上的吞吐量约为同类开源模型的 6 倍,同时保持相同准确率。
原文给出 550B 开源模型的权重、训练数据与完整配方,并说明其在长任务智能体上的吞吐表现,读者可据此判断开源前沿模型的可复现程度。
1/ NVIDIA shipped Nemotron 3 Ultra today, a fully open 550B model with 55B active params, with the weights, training data, and complete recipe all released openly. That alone is rare at this scale.
The headline however actually is speed. Ultra is a hybrid Mamba-Attention MoE, an architecture built for fast decoding and a light memory footprint over long contexts, and NVIDIA clocks it at roughly 6x (!) the throughput of comparable open models on long-output agent workloads while holding the same accuracy.
That's a serious engineering result, and it's aimed exactly where the industry is heading: autonomous agents that run long, multi-turn tasks where throughput per GPU is what actually costs money.
It was pre-trained in 4-bit (NVFP4) across 20T tokens, the largest stable run of its kind shown to date. And the post-training introduces MOPD, where ten-plus specialist teacher models distill their skills into the student on its own rollouts, sometimes pushing it past the teachers themselves.
The interesting aspect:This is a frontier-class model you can fully reproduce.
Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. Video在 X 查看被引用的帖子
来源:@kimmonismus · x.com