跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 23 天前AI 评分32
AI 导读

SemiAnalysis 指出 HiCache 作为 per-rank 主机侧 L2 缓冲位于 HBM 与外部 DRAM KV store 之间,每次加载和卸载都多一次拷贝,L3 到 L2 的取数未与计算流水线化,且在 MLA+TP 下每个 rank 会通过 PCIe 和主机内存复制同一份 KV,仅凭本地信息做驱逐决策。

正文

HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is not pipelined with compute, and under MLA with TP each rank replicates the same KV across PCIe and host memory while making eviction decisions on local information only.

UMBP through the SGLang KVCache Store Linker replaces that with a direct, layer-wise pipelined HBM-to-DRAM path into a single shared per-node pool, which removes the intermediate copy and stall, deduplicates the replicated KV, and lets placement and eviction be decided globally across DP ranks and instances. (3/3)

来源:@SemiAnalysis_ · x.com