跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-04AI 评分38
AI 导读

GPT-6 Astra 在 Perplexity 的 WANDR 基准上取得 0.682 分,为所有测试模型最高,单任务成本 $11.98。它比 Fable 5.1 高 13.5% 且成本低 6.1%,比 Opus 5 高 27.0% 但成本高 3.3%。WANDR 含 500 个公开任务、170,495 条来源记录,同时考核精确率与完成度,硬分要求整条分支全对才满分。

正文

GPT-6 Astra delivered Perplexity’s strongest WANDR result yet

13.5% above Fable 5.1 while costing 6.1% less per task.

This benchmark WANDR is quite unusual because it tests wide-and-deep research capability of models: finding large sets of qualifying entities, then backing every requested fact with checkable evidence.

WANDR itself contains 500 public tasks requiring 170,495 source-backed records, so incomplete research is directly penalized even when the facts an agent did find are correct.

Its scoring tracks both precision and completion, with stricter hard scores requiring an entire requested branch to be correct before receiving full credit.

So the 0.682 result points to a substantial gain on long, evidence-heavy research work where an agent must keep finding, checking, and organizing information at scale.

引用@perplexity_ai@perplexity_ai
We evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost. https://t.co/SyYmD38qvq
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com