跳到正文
elvis· @omarsar0 · X·· 2 小时前AI 评分50
AI 导读

一项名为SelfSearch的工作让编码智能体改写自身harness,用DeepSeek V4 Flash在Terminal-Bench 2.1上解决82.0%任务,搜索过程无任务奖励,成本仅4.03美元,追平公开九harness对比中的Codex。

正文

Learn to optimize your own harness, folks.

You can squeeze much more performance and value from a harness.

This work claims that a self-modified harness matched Codex for $4.03.

Specifically, a coding agent rewrote its own harness until it solved 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, without any task reward during the search.

The search cost was $4.03.

That matches Codex, the top harness in a public nine-harness comparison run under the same settings.

SelfSearch has agents modify their own instructions, tools, and procedures using records of earlier self-modification attempts. Each record holds the reasoning, tool actions, and outcomes. The modified agent then becomes the next improver.

Population-mean success rises in all six model-benchmark settings, with single agents gaining up to 11.2 points. On SWE-bench Multilingual, one evolved agent gains 5.0 points and spends 38.5% less on tasks both versions solve.

Paper: https://arxiv.org/abs/2609.37968

Chat with Paper: https://academy.dair.ai/papers/selfsearch-reward-free-search-for-self-improving-agents-2609.37968

来源:elvis · x.com