Anthropic 公开三项衡量 AI 研发进展的指标,截至 2026 年 8 月,Claude 主导其 26% 的模型研发工作,90% 以上的工作处于 AI 协作或更高层级。
Anthropic 公开内部 AI 研发自动化测量口径与月度数据,读者可据此看到智能体参与研发的实际比例。
Inside Anthropic 26% of its AI R&D work has reached the point where Claude can do most of the task end-to-end from a high-level instruction.
A human mainly supervises.
- “As of Aug 2026, there were approximately 30,000 agents doing research and engineering work at Anthropic at any one time in our most-used internal platform.”
- More than 90% of measured Anthropic R&D already has AI doing substantial chunks of the work or more
“The share of work at or above “AI collaborates” is above 90%.”
- More than a billion agent decisions in one month is the scale here; the online monitor blocked roughly one in every 47,000.
“We analyzed over a billion decisions from our research & engineering agents over August 2026 and found that 0.002% of them (about 1 in 47,000) were blocked by this monitor.”
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public. Today, we're sharing three measurements that help track AI development: 1. How much AI R&D is done by AI. 2. How well AI agents are overseen. 3. How compute is allocated. We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them. As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information. Read the full post and methodology: https://t.co/iPFz8Z4ugE在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com