跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 17 天前AI 评分36
AI 导读

ECDYSIS 提出按跨任务的重复失败模式而非失败次数来训练 LLM Agent 运行时,避免为每次错误改写运行时而学错教训。该方法将不同任务中的同类失败归组后再决定是否修复 harness,带来更高准确率、更快训练和更强的跨模型迁移。

正文

If an agent rewrites its runtime for every bad decision, it can learn the wrong lesson

Fix recurring failures across tasks, not every failure.

Agent runtimes should learn from failure patterns, not failure counts: ECDYSIS grouped recurring problems across tasks and delivered higher accuracy, faster training, and stronger cross-model transfer.

Patch every miss, and you can accidentally hard-code one model's bad habits into the system.

ECDYSIS instead looks for the same kind of failure across different tasks before deciding the harness itself needs fixing.

– arxiv. org/abs/2609.11677

Title: "Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents"

来源:@rohanpaul_ai · x.com