AI 导读
若这40万块GPU为NVIDIA Rubin,按其公布的NVFP4规格,理论峰值训练性能约为10万块Blackwell GPU的14倍,推理性能约为20倍。拆解来看:GPU数量为4倍、单卡训练性能3.5倍、单卡推理性能5倍,纸面等效于约140万块Blackwell用于训练、200万块用于推理。
正文
If those 400,000 GPUs are Rubin, NVIDIA’s published NVFP4 specs imply roughly 14× the theoretical peak training performance and 20× the theoretical peak inference performance of 100,000 Blackwell GPUs.
-4× as many GPUs.
-3.5× training performance per GPU.
-5× inference performance per GPU.
That’s the equivalent of roughly 1.4 million Blackwell GPUs for training, or 2 million for inference, on paper.
From a hardware perspective, there are probably no more obstacles to the automated searcher. And RSI is truly within reach.
NVIDIA long.
来源:@kimmonismus · x.com