跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-24AI 评分44
AI 导读

这篇论文提出用主动推理让 AI 智能体显式决定每次澄清、检索或工具调用是否值得消耗 token 与延迟。在受控测试中,前沿模型能逐步缩小隐藏答案范围,但提问次数仍多于最优规划器;生成任务中定向澄清将验证器合规率从 0.0417 提升至 0.375,平均 token 用量从约 112 增至 219。

正文

AI agents have an expensive problem: when a request is missing information, they either guess too early or keep asking, retrieving, and calling tools without knowing whether more context is worth the cost.

This paper uses active inference to make that choice explicit: every clarification, retrieval, tool call, or prompt test should earn its tokens, latency, or user effort by reducing uncertainty that matters to the final answer.

In controlled tests, frontier models steadily narrowed down hidden answers, yet still used more questions than an optimal planner.

On a generation task, targeted clarification raised verifier compliance from 0.0417 to 0.375, while average token use rose from about 112 to 219.

The practical recommendation is to give agents an explicit context-acquisition layer instead of treating context gathering as an automatic behavior.

Before each extra step, decide whether to ask, retrieve, inspect, or act.

Get only the missing context that changes the decision, then stop.

来源:@rohanpaul_ai · x.com