跳到正文
@omarsar0· @omarsar0 · X·· 15 天前AI 评分46
AI 导读

论文提出 Question's Gambit,在智能体开始搜索前先运行一次:把问题拆成线索、各自转为互补搜索、汇总结果并重排,让智能体带着已排序的证据集进入循环。

正文

Impressive paper showing how much the first retrieval step matters for deep research agents.

It helps to improve GPT-5.5 from 83.1% to 90.5% on BrowseComp-Plus with the same retriever and the same agent loop.

It seems that the gain comes from the opening context.

The authors propose Question's Gambit which runs once, before the agent starts searching.

It splits the question into clues, turns each clue into complementary searches, pools the results, and reranks them.

The agent then starts its loop with that ranked set already in context.

The same change lifts GPT-5.4-mini from 68.1% to 79.0% and DeepSeek-v4-pro from 71.4% to 76.9%, and roughly halves calibration error for GPT-5.5. It costs between 2.3 and 5.3 extra tool calls per question.

In an error analysis, only 3 of the 79 remaining GPT-5.5 errors come from the gold document never being retrieved. The other 76 happen later, when the agent previews, opens or uses the evidence.

Paper: https://t.co/MBUSbcsEfV

Chat with Paper: https://t.co/tNAuFse1k6

来源:@omarsar0 · x.com