跳到正文
@perplexity_ai· @perplexity_ai · X·· 2026-09-05AI 评分24
AI 导读

Tulip 是 Perplexity 的轻量级 Rust gRPC 推理服务器,位于 Ivy 和 ROSE 引擎之间。它收集传入请求,将其批处理,然后发送到 GPU。 对于小型嵌入向量模型,运行时间取决于 token 数,而非查询数量,因此约 512 个 token 就能填满 GPU。https://t.co/2hTVDKIndU

正文

Tulip is Perplexity’s lightweight Rust gRPC inference server that sits between Ivy and the ROSE engine. It collects incoming requests, batches them, and sends them to the GPU.

For small embedding models, runtime depends on tokens, not query count, so ~512 tokens fills the GPU. https://t.co/2hTVDKIndU

来源:@perplexity_ai · x.com