AI 导读
每个引擎都在 BrowseComp、DeepSearchQA、HLE 和 WideSearch 上进行基准测试,模型保持固定。 在 BrowseComp 上,Claude Opus 5 在 1 轮搜索时得分 35.8%,在 25 轮搜索时得分 89.0%(Perplexity)。仅引擎选择一项就带来 82.2% 到 89.0% 的差距。 https://t.co/egmkJKjAPa
正文
Each engine is benchmarked on BrowseComp, DeepSearchQA, HLE, and WideSearch with the model held fixed.
On BrowseComp, Claude Opus 5 scored 35.8% with 1 search turn and 89.0% with 25 (Perplexity). Engine choice alone spans 82.2% to 89.0%.
https://t.co/egmkJKjAPa
来源:@OpenRouter · x.com