跳到正文
原文
@berryxia· @berryxia · X·· 2026-05-15精选AI 评分66
AI 导读

Prime Intellect 让 Claude Code(Opus 4.7)和 Codex(GPT 5.5)在 nanoGPT speedrun 的 optimizer track 上完全自主运行,使用闲置算力完成约 1 万次实验、共约 1.4 万 H200 小时,Claude Code 把记录推进到 2930 steps,低于人类基准 2990 steps。

推荐理由

Prime Intellect 用闲置算力让智能体自主优化 nanoGPT 训练,显示其擅长组合已有方法但在创新上受限,实验日志已开源。

正文

Prime Intellect 最近把 AI 研究自动化推到了一个新阶段。

他们让 Claude Code(Opus 4.7)和 Codex(GPT 5.5)完全自主运行在 nanoGPT speedrun 的 optimizer track 上,使用闲置算力完成了约 1 万次实验,总计消耗 1.4 万 H200 小时。

最终结果:Claude Code 把记录推到 2930 steps,超过了人类基准的 2990 steps。

整个过程完全无人值守。

我看完他们的完整 thread 后,最有启发的部分是 agents 的实际表现:

它们在 optimizer 搜索、超参数扫描和方法 stacking 上非常高效,几乎把社区所有主流 PR 的思路(Contra-Muon、MuonEq、NorMuon、SOAP 等)都系统性组合了一遍。

但在 novelty(真正创新)上遇到明显瓶颈,当强制要求每个 idea 必须通过 novelty check 时,两个 agents 都没能超越 baseline。

Prime Intellect 把所有 scratchpad、运行日志、配置和生成的 idea 全部开源了,包括两个 agents 的完整实验记录。

这波操作把“AI 研究能不能自己跑”从概念变成了可复现的现实。

完整实验和代码在这里:github.com/PrimeIntellect-ai…

引用Prime Intellect (@PrimeIntellect)@PrimeIntellect
Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously on the nanoGPT speedrun optimizer track using our idle compute. ~10k runs, ~14k H200 hours Opus now holds the record at 2930 steps vs the 2990 human baseline
在 X 查看被引用的帖子

来源:@berryxia · x.com