跳到正文
DAIR.AI· @dair_ai · X·· 2 天前AI 评分52
AI 导读

NVIDIA 发布 Mid-Harness 论文,提出在生成器与 harness 之间采样多个候选动作并先验证再执行。TerminalBench-Lite 上用 GPT-5.6 Sol 验证器从 8 个采样动作中选择,Pass@1 从 50.0% 升到 68.0%;弱验证器下多采样几乎无收益。

正文

Banger paper from NVIDIA on test-time compute for terminal agents.

The finding is that you should sample several candidate shell commands, verify them before running one, and spend more on the verifier than on extra samples.

With a GPT-5.6 Sol verifier choosing among 8 sampled actions, TerminalBench-Lite Pass@1 rises from 50.0% to 68.0%. With a weak verifier, extra samples add almost nothing.

Mid-Harness leaves the generator and harness unchanged and works between them. When a small TMAX-9B model verifies its own candidates, pairwise comparison works best, and distilling the strong verifier into it helps further.

Combining action sampling with trajectory sampling reaches higher success at lower estimated token cost than sampling full trajectories alone.

Paper: https://academy.dair.ai/papers/mid-harness-scaling-actions-between-model-and-harness-for-terminal-agents-2609.39982

来源:DAIR.AI · x.com