AI 导读
评估 LLM 推理的一个有用指标是 tok/s/MW:系统每消耗一兆瓦全部配置电力每秒能生成多少 token。 在单并发用户下,下图的 B300 配置在 DeepSeek V4 上每 GPU 生成约 14 output tok/s。SemiAnalysis 数据中心行业模型估算每块 B300 配置电力为 1.9 kW(0.0019 MW)。 14 tok/s ÷ 0.0019 MW = 7,368 tok/s/MW 即每配置一兆瓦约 7,400 output tok/s。(1/4)🧵
正文
A useful metric for evaluating LLM inference is tok/s/MW: how many tokens a system generates per second for each megawatt of all-in provisioned power.
At one concurrent user, the B300 configuration below generates ~14 output tok/s per GPU on DeepSeek V4. The SemiAnalysis Datacenter Industry Model estimates 1.9 kW (0.0019 MW) of provisioned power per B300.
14 tok/s ÷ 0.0019 MW = 7,368 tok/s/MW
That’s roughly 7,400 output tok/s per provisioned megawatt. (1/4)🧵
来源:@SemiAnalysis_ · x.com