跳到正文
@kimmonismus· @kimmonismus · X·· 29 天前AI 评分29
AI 导读

若这40万块GPU为NVIDIA Rubin,按其公布的NVFP4规格,理论峰值训练性能约为10万块Blackwell GPU的14倍,推理性能约为20倍。拆解来看:GPU数量为4倍、单卡训练性能3.5倍、单卡推理性能5倍,纸面等效于约140万块Blackwell用于训练、200万块用于推理。

正文

If those 400,000 GPUs are Rubin, NVIDIA’s published NVFP4 specs imply roughly 14× the theoretical peak training performance and 20× the theoretical peak inference performance of 100,000 Blackwell GPUs.

-4× as many GPUs.
-3.5× training performance per GPU.
-5× inference performance per GPU.

That’s the equivalent of roughly 1.4 million Blackwell GPUs for training, or 2 million for inference, on paper.

From a hardware perspective, there are probably no more obstacles to the automated searcher. And RSI is truly within reach.

NVIDIA long.

来源:@kimmonismus · x.com