AI 导读
Lily 将 prefill 和 decode 视为本质上不同的工作负载。 Prefill 一次性处理大量提示词 token,可以在它们之间复用权重,而 decode 一次只生成一个 token,复用少得多,因此内存流量和带宽更为重要。https://t.co/rhr5pSGBZG
正文
Lily treats prefill and decode as fundamentally different workloads.
Prefill processes many prompt tokens at once and can reuse weights across them, while decode generates one token at a time with much less reuse, making memory traffic and bandwidth more important. https://t.co/rhr5pSGBZG
来源:@perplexity_ai · x.com