跳到正文
@EpochAIResearch· @EpochAIResearch · X·· 2026-08-27AI 评分30
AI 导读

Gemma 4 的 K=V 优化使其理论 INT8 KV 缓存降至每 100K 上下文请求约 2.5 GB,在扣除约 31 GB 的 FP8 权重后,128 GB SRAM 中最多可常驻约 39 个请求,尚未计入运行时开销。Nvidia 未披露 LPX 实际的 KV 布局。

正文

Gemma 4’s K=V optimization brings its theoretical INT8 KV cache to ~2.5 GB per 100K-context request, implying up to ~39 resident requests in 128 GB of SRAM after ~31 GB of FP8 weights, before runtime overhead. Nvidia has not disclosed the actual LPX KV layout.

来源:@EpochAIResearch · x.com