Anthropic 发布研究论文,让 Claude 承担改进其他 AI 模型对齐的任务,结果 10 项被测对齐缺陷全部改善,且未降低已测的通用能力。
材料呈现了 AI 研究 AI 的闭环雏形,读者可据此了解自动化对齐研究目前走到了哪一步。
Anthropic just published one of the clearest previews of recursive self-improvement yet. But we still dont talk about it.
Yesterday Anthropic released a reserach paper. Claude was tasked with improving the alignment of other AI models. It searched the literature, proposed methods, created training data, trained the models, evaluated the results and iterated.
It improved all 10 tested alignment failures without degrading measured general capabilities. Anthropic even used the weaker Sonnet 5 to post-train an early Opus 4.8 checkpoint, bringing its alignment close to the released production model within 60 hours.
Anthropic:
“The best AAR method beats what experienced humans propose, on average within six hours. (...) Human-guided research directions do not lead to stronger performance.”
This is not full recursive self-improvement yet. The improved model did not become the next researcher and repeat the process. But most of the loop now exists:
AI researches AI.
AI trains improved AI.
AI evaluates the result.
AI iterates.
The loop closes when the improved AI becomes the researcher for the next generation. That is when progress could begin to compound. And tbh Id say we are pretty close to it. So talk that reserach serious.
h/t to @tradernewsai for bringing this to my attention
Sources: Anthropic blog / Tech Crunch
来源:@kimmonismus · x.com