跳到正文
@AravSrinivas· @AravSrinivas · X·· 2026-09-05AI 评分34
AI 导读

深入解析 Perplexity 如何大规模提供搜索结果:用嵌入向量做排序、基于 GPU 的模型推理、请求批处理、运行推理服务器,以及处理延迟/吞吐量的权衡。https://t.co/2Qy8HhHRVQ

正文

A deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based model inference, request batching, running inference servers, and handling latency/throughput trade-offs. https://t.co/2Qy8HhHRVQ

引用@perplexity_ai@perplexity_ai
Every answer in Perplexity starts with embedding and ranking models picking the most relevant results for the query. Today we published research on how we built SoTA serving infrastructure behind those models. Read the research: https://t.co/6Rt2XlxoXA https://t.co/ORGt9TjkR0
在 X 查看被引用的帖子

来源:@AravSrinivas · x.com