跳到正文
@omarsar0· @omarsar0 · X·· 2026-09-02AI 评分47
AI 导读

字节跳动Seed提出HarnessDev,不再用模型完成的任务打分,而是评估其构建的agent harness。方法从可运行的弱种子和少量案例出发构建完整执行系统,第二阶段再交回harness要求其根据下游反馈改进,两阶段均按能力与执行token成本评分。

正文

Banger paper from ByteDance Seed.

If you are curious about self-evolving agent harnesses, this one is worth your time.

(bookmark it)

The proposes method, HarnessDev, stops scoring a model on the tasks it completes and scores it on the harness it builds.

The agent starts from a weak but runnable seed plus a handful of cases, then builds a full execution system.

A second stage hands that harness back and asks it to improve on downstream feedback.

Both stages are scored on capability and on execution-token cost, so there is awareness of efficiency and spend.

The experiments include six creator LLMs, four domains, 2,207 held-out downstream instances.

The result splits by domain.

Generated harnesses stay well behind mature human-engineered references on code and on search and research, while matching or beating them on writing and machine-learning experimentation.

Evolution produces gains, but they are unstable and transfer only partially to held-out tasks, and they depend heavily on which model runs the harness.

Paper: https://t.co/nbXaLrXU3o

Chat with Paper: https://t.co/ithbgySYDU

来源:@omarsar0 · x.com