AI 导读
倒数排名融合(RRF)子采样仅用 31.7% 的文档,就能在 Combined Recall@1000 上保留全语料库的模型排名。 对于 pplx-embed-v1-4b,这将评测从 4,608 H200 GPU-hours 降至约 1,500 H200 GPU-hours。https://t.co/ZPTvJ1xYoN
正文
Reciprocal rank fusion (RRF) subsampling preserves the full-corpus model ranking on Combined Recall@1000 using only 31.7% of the documents.
For pplx-embed-v1-4b, this reduces evaluation from 4,608 to roughly 1,500 H200 GPU-hours. https://t.co/ZPTvJ1xYoN
来源:@perplexity_ai · x.com