AI 导读
Groq 3 LPX 机架共有 128 GB 的 SRAM(比单块 288 GB HBM 的 B300 GPU 还少),但速度快得多,达到 40 PB/s 对 8 TB/s(约 5,000 倍)。这使其非常适合对权重和 KV cache 能装进其 SRAM 容量的小模型进行超高速解码。在长上下文长度下,KV cache 会占用大量内存。
正文
The Groq 3 LPX rack has a total of 128 GB of SRAM (less than a single B300 GPU with 288 GB of HBM), but is much faster at 40 PB/s vs. 8 TB/s (~5,000×). This makes it very good for ultrafast decode of small models whose weights and KV cache fit within its SRAM capacity. At long context lengths, the KV cache can use a lot of memory.
来源:@EpochAIResearch · x.com