字节跳动Seed提出HarnessDev,不再用模型完成的任务打分,而是评估其构建的agent harness。方法从可运行的弱种子和少量案例出发构建完整执行系统,第二阶段再交回harness要求其根据下游反馈改进,两阶段均按能力与执行token成本评分。
Banger paper from ByteDance Seed.
If you are curious about self-evolving agent harnesses, this one is worth your time.
(bookmark it)
The proposes method, HarnessDev, stops scoring a model on the tasks it completes and scores it on the harness it builds.
The agent starts from a weak but runnable seed plus a handful of cases, then builds a full execution system.
A second stage hands that harness back and asks it to improve on downstream feedback.
Both stages are scored on capability and on execution-token cost, so there is awareness of efficiency and spend.
The experiments include six creator LLMs, four domains, 2,207 held-out downstream instances.
The result splits by domain.
Generated harnesses stay well behind mature human-engineered references on code and on search and research, while matching or beating them on writing and machine-learning experimentation.
Evolution produces gains, but they are unstable and transfer only partially to held-out tasks, and they depend heavily on which model runs the harness.
Paper: https://t.co/nbXaLrXU3o
Chat with Paper: https://t.co/ithbgySYDU
来源:@omarsar0 · x.com