跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-31AI 评分45
AI 导读

字节跳动新论文提出 Chain-of-Experience:在测试时把此前的尝试与反馈保留在上下文中让模型重试,效果优于将历史压缩成整洁的记忆摘要。在 6 个数学、代码与知识基准上,自反馈平均达 71.0%,无反馈的迭代求解为 66.8%,正确性或执行器反馈达 79.3%。

正文

New ByteDance paper shows for test-time improvement, keeping the messy history of attempts can work better than turning that history into a neat memory summary.

Chain-of-Experience keeps earlier attempts and feedback in context, then asks the model to try again. Across 6 math, coding, and knowledge benchmarks, self-feedback averaged 71.0%, versus 66.8% with iterative solving but no feedback; correctness or executor feedback reached 79.3%.

The paper also reports 5.6% overall improvement with 19% lower API cost across tasks and models when feedback is used.

No weights change here, so this is contextual adaptation rather than persistent learning. Self-feedback also hurt on BrowseComp-Plus when solving the task required external search.

For agents, preserve the trajectory, add reliable feedback, and compress only when you know what can safely disappear.

来源:@rohanpaul_ai · x.com