跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-29AI 评分57
AI 导读

Anthropic Fellows 的研究显示,Claude 可在 48 小时和 1 张 GPU 的条件下自主提出对齐训练方法并训练、测试小模型,其效果超过 28 位经验研究者的一次性方案。研究让 5 个智能体并行工作最多 48 小时并共享结果,同时用隐藏测试和能力门控过滤过拟合与明显的能力退化。发帖作者补充称,被纠正的模型能力比执行纠正的模型更强。

正文

New Anthropic research shows Claude aligning Claude all by itself.

And here the model being corrected was more capable than the model correcting it

Looks like another great research at the frontier of recursive self-improvement. AI is reaching the point where it can do much of the research required to make its own stronger successors safer.

The loop it describes closes on itself: propose, train, score, repeat, no researcher required.

Claude autonomously discovered safety-training methods that outperformed one-shot ideas from 28 experienced researchers.

Five agents worked in parallel for up to 48 hours, sharing results while hidden tests and capability gates screened out overfitting and obvious regressions.

引用@AnthropicAI@AnthropicAI
New Fellows Research: Can Claude autonomously align other AIs? We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well. https://t.co/nhlCMgQl46
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com