跳到正文
@perplexity_ai· @perplexity_ai · X·· 2026-09-03AI 评分21
AI 导读

Lily 将 prefill 和 decode 视为本质上不同的工作负载。 Prefill 一次性处理大量提示词 token,可以在它们之间复用权重,而 decode 一次只生成一个 token,复用少得多,因此内存流量和带宽更为重要。https://t.co/rhr5pSGBZG

正文

Lily treats prefill and decode as fundamentally different workloads.

Prefill processes many prompt tokens at once and can reuse weights across them, while decode generates one token at a time with much less reuse, making memory traffic and bandwidth more important. https://t.co/rhr5pSGBZG

来源:@perplexity_ai · x.com