Claude 5 Fable 在几乎所有已测试的 AI 能力基准上达到 SOTA,在软件工程、知识工作、视觉与科研方面表现突出。作者称其比过往 Claude 模型更省 token,能在数百万 token 的长任务中保持专注并用自记笔记改进输出。Stripe 早期测试称该模型把数月工程压缩到数天,在 5000 万行 Ruby 代码库中用一天完成原本需团队两个多月的全库迁移。
原文汇总了 Fable 5 的基准成绩和 token 效率变化,并借 Stripe 的迁移案例展示其长任务表现。
Claude 5 Fable tl;dr
- It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research
-The longer and more complex the task, the larger Fable 5’s lead over our other models
-its more token-efficient than past Claude models
- Fable 5 stays focused across millions of tokens in long-running tasks and improves its outputs using its own notes
Fable 5 is more than just better benchmarks. It's more efficient, allows for longer work periods, offers better context management, and so much more.
GPT-5.6 is just around the corner.
I'm a huge Codex fan, but Fable/Mythos is in a league of its own. I'm curious to see if OpenAI will release its own Mythos.
"During early testing, Stripe reported that Fable 5 compressed months of engineering into days. In a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand."
Claude 5 Fable Benchmarks! Holy moly, significant jump even to Mythos在 X 查看被引用的帖子
来源:@kimmonismus · x.com