跳到正文
@perplexity_ai· @perplexity_ai · X·· 2026-09-05AI 评分31
AI 导读

ROSE 是模型引擎。 它为 LLM 和嵌入向量复用相同的 kernel。对于嵌入向量,它跳过 KV cache,并使用 ragged attention 而非 paged attention。 ROSE 支持多种注意力后端,因此 kernel 的选择取决于模型形态和序列长度。https://t.co/cdBEP3m3wG

正文

ROSE is the model engine.

It reuses the same kernels for LLMs and embeddings. For embeddings, it skips the KV cache and uses ragged attention instead of paged attention.

ROSE supports multiple attention backends, so kernel choice depends on model shape and sequence length. https://t.co/cdBEP3m3wG

来源:@perplexity_ai · x.com