跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-24AI 评分43
AI 导读

AutoDesign 能自动运行智能体任务、定位失败原因,并改写提示词、工具、验证检查或重试规则等脚手架组件,改动只有在训练任务提分且不损害留出集时才保留。在论文转会议海报任务上,7 个智能体各提升 5 至 19.6 分,最便宜的模型涨幅最大,团队同时发布了评测基准 PosterBench。结论是脚手架对弱模型的收益最高,升级更贵模型前不妨先花一周改进现有模型的检查与重试逻辑。

正文

A weak model with well-built scaffolding around it can close most of the gap to a frontier model, so fix your code before you upgrade your model.

Scaffolding pays off in inverse proportion to model strength, so the money you save by switching to a cheap model can be recovered by engineering around it.

They built a system called AutoDesign that does this automatically. It runs an agent on real tasks, looks at what went wrong, then rewrites one piece of the surrounding setup: a prompt, a tool, a validation check, a retry rule.

A change survives only if it improves scores on the training tasks and doesn't hurt a held-out set, so the system can't just overfit its way upward.

They tested it on turning papers into conference posters, and released a benchmark, PosterBench, to score them.

The result is that scaffolding is worth more than most people assume, and worth the most to weak models. Seven agents each gained 5 to 19.6 points, with the biggest jumps going to the cheapest models.

So before you upgrade to a pricier model, spend a week improving the checks and retry logic around the one you have.

– arxiv. org/abs/2608.13560

Title: "AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design"

来源:@rohanpaul_ai · x.com