AI 导读
Perplexity 将查询和文档嵌入到同一个向量空间,然后通过最近向量进行搜索。 这带来了两种工作负载:用于索引和评分的批量嵌入(侧重吞吐量),以及用于实时搜索的逐查询在线嵌入(侧重延迟)。
正文
Perplexity embeds queries and documents into one vector space, then searches by nearest vectors.
This creates two workloads: bulk batch embedding for indexing and scoring (throughput‑focused) and per‑query online embedding for live search (latency‑focused).
来源:@perplexity_ai · x.com