@swyx· @swyx · X·· 2026-06-10精选AI 评分70
AI 导读
karpathy 称 Claude Fable 5 与 Mythos 是同一底层模型,只是增加了防护措施,在各项基准上都以一定优势达到 SOTA,并认为这是值得大版本号跃迁的阶跃式进步,尤其擅长在极难问题上的长时间解题会话。swyx 转发表示自己重跑了历史图表上的 FC Diamond,认为官方表格和图表都没有体现出这种起飞幅度,因为 Fable 属于不同级别的模型。
推荐理由
转发的评测者认为官方榜单未体现 Fable 5 的进步幅度,可作为判断这次模型跃迁的定性参考。
正文
just finished rerunning FC Diamond on my historical charts. none of the official tables/charts are capturing the degree of takeoff.
nitter.net/karpathy/status/206440…
its this same chart all the way down difficulty classes (below) breaks every curve fit because Fable is a diffferent CLASS of model, with beeeeeg model smell.
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time. I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!在 X 查看被引用的帖子
来源:@swyx · x.com