AI 导读
ROSE 是模型引擎。 它为 LLM 和嵌入向量复用相同的 kernel。对于嵌入向量,它跳过 KV cache,并使用 ragged attention 而非 paged attention。 ROSE 支持多种注意力后端,因此 kernel 的选择取决于模型形态和序列长度。https://t.co/cdBEP3m3wG
正文
ROSE is the model engine.
It reuses the same kernels for LLMs and embeddings. For embeddings, it skips the KV cache and uses ragged attention instead of paged attention.
ROSE supports multiple attention backends, so kernel choice depends on model shape and sequence length. https://t.co/cdBEP3m3wG
来源:@perplexity_ai · x.com