Rich Sutton 谈 AI 创造力与发现
强化学习先驱 Rich Sutton 就 AI 创造力与发现发表观点,相关讨论在 Hacker News 上获得 114 分、56 条评论。正文未提供更多细节。
关注 AI 研究者、开发者与机构的动态
强化学习先驱 Rich Sutton 就 AI 创造力与发现发表观点,相关讨论在 Hacker News 上获得 114 分、56 条评论。正文未提供更多细节。
因为害怕额度问题,所以先让他找问题、写计划,没让他改代码。 结果发现他找问题找得老快了,而且也挺准的
在我 26 万行代码的 CodePilot 代码库中尝试 Fable 5,看一下它能找出多少问题
We've reset usage limits across our products! For those just starting to test Fable, here's four tips for using it more effectively: 1. Give it bigger, more ambitious tasks than what previous models could handle. 2. Use xhigh/high effort as your default for best performance, med for faster interactive sessions. 3. Rework your skills and CLAUDE.mds. Instructions written for prior models anchor Fable to stale patterns, let it use its own judgment first. 4. Move from providing tasks to providing objectives. Describe what done looks like and how to verify it, then let Fable find the path (/loop and /goal are built for this)
Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.
推荐理由:借 Anthropic 与 OpenAI 的模型命名差异,作者给出一个观察两家产品调性的轻量切口。
猴哥都主动祝贺Anthropic 的Fable 5 了。 那么,问题来了。 Google 人呢? 虽然,Google 也是A社30 亿美金的大股东,但也要努力啊!
congrats to the Anthropic team on Fable!!
让他发挥…… @LumaLabsAI Ray3.2 现已上线 → lumalabs.ai/ray3-2 视频
有用户发现购买高铁票时可直接选择静音车厢,车厢内没有手机外放短视频的噪音,乘务员还会当场制止外放行为,旅途 Vibe Coding 体验极佳。
6.22 日 后只能调用API使用了! 到时候是不是就知道真正的“中转站”到底是不是真的有“中转”模型Fable5 了😂
Fable 5 的这个“Included until June 22”是什么意思?订阅用户只在六月22号前才能免费体验一下么,后面要单独收费吗?
just finished rerunning FC Diamond on my historical charts. none of the official tables/charts are capturing the degree of takeoff. nitter.net/karpathy/status/206440… its this same chart all the way down difficulty classes (below) breaks every curve fit because Fable is a diffferent CLASS of model, with beeeeeg model smell.
Gemini 3.5 Live Translate is now in Public Preview via the Gemini API, delivering low-latency speech-to-speech translation across 70+ languages and 2,000 language pairs! 🌍 Challenge time: What is the most niche, unique, or complex language pair your application needs to translate? Tell us in the comments, and we’ll let you know if the model supports it! Full blog link: goo.gle/3QzaHwN Video
推荐理由:Google 将 Gemini 3.5 实时语音翻译推向公开预览,可了解其对小众语言对的覆盖范围与 API 接入方式。
x.com/i/article/206447998310…
Google 开发者账号宣布,Gemini 模型现已通过 Apple 的 Foundation Models 框架和 Xcode 中的原生支持,向数百万 Apple 开发者开放。
Gemini models are now accessible to millions of Apple developers through Apple’s Foundation Models framework and natively within Xcode. You can now easily swap between local and cloud inference using a shared API surface to build next-generation agentic app experiences, increase development velocity, and offload heavy workloads to the cloud. Additionally, you can use agentic coding assistance from Gemini in Xcode to accelerate multi-step development tasks. Check out the full announcement to get started: goo.gle/3Q1YDnD Video
推荐理由:该消息交代了 Gemini 接入 Xcode 与 Foundation Models 框架的具体入口,便于判断 Apple 开发者的模型选择格局。
Cohere 以 Apache 2.0 许可开源 North Mini Code,这是一款总参数 30B、激活参数 3B 的 MoE 小模型,主打 agentic coding。
Small: 30 billion parameters, 3B active. Efficient: Benchmarks to 33.4 on the Artificial Analysis Coding Index, competitive among similar sized models. Open Source: Apache 2.0 license so developers can experiment, test, and build their way. Learn more: cohere.com/blog/north-mini-c‌
这也能看出 Fable5 的效果有点明显啊,交互细节和动效都很到位。
We've reset usage limits across our products! For those just starting to test Fable, here's four tips for using it more effectively: 1. Give it bigger, more ambitious tasks than what previous models could handle. 2. Use xhigh/high effort as your default for best performance, med for faster interactive sessions. 3. Rework your skills and CLAUDE.mds. Instructions written for prior models anchor Fable to stale patterns, let it use its own judgment first. 4. Move from providing tasks to providing objectives. Describe what done looks like and how to verify it, then let Fable find the path (/loop and /goal are built for this)
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. Video
推荐理由:原文列出该模型基准几乎全线 SOTA,并说明敏感领域会自动回退到 Opus 4.8 的机制。
Claude 5 Fable tl;dr - It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research -The longer and more complex the task, the larger Fable 5’s lead over our other models -its more token-efficient than past Claude models - Fable 5 stays focused across millions of tokens in long-running tasks and improves its outputs using its own notes Fable 5 is more than just better benchmarks. It's more efficient, allows for longer work periods, offers better context management, and so much more. GPT-5.6 is just around the corner. I'm a huge Codex fan, but Fable/Mythos is in a league of its own. I'm curious to see if OpenAI will release its own Mythos. "During early testing, Stripe reported that Fable 5 compressed months of engineering into days. In a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand."
OpenRouter 是 AgentCard 上购买量第三的服务,仅次于 Amazon 和 OpenAI ❤️
591 agents went shopping. Nobody supervised. 3 hours & thousands of $$$ later, here's what they bought: 1. @amazon 2. @OpenAI 3. @OpenRouter 4. @DoorDash 5. @AnthropicAI 6. @PayPal @agentcardai
有讨论指出,Mythos/Fable 会刻意阻碍涉及 AI 研究与开发(AI Research Development)的请求。该帖在 Hacker News 获得 7 分、1 条评论,正文未提供更多技术细节或官方回应。


The guardrails are way too strict. Even the simplest questions get cut off immediately. And it's only on the schedule until June 22nd. Damn, Anthropic really thinks the model is too powerful.


Direction goes in. Cinema comes out. Ray3.2 is here → lumalabs.ai/ray3-2 Video
我理解 Anthropic 对模型在没有护栏的情况下被滥用的担忧很重要。我也认真对待这一点。我们谈论的是一项具有未知潜力的技术。 然而,它在某些情况下确实完全无法使用,这一点令人遗憾。
Claude Fable 5 is unusable at this time. How the hell is this prompt a cybersecurity or biology risk?! Almost every prompt I've tried gives me the same error! What’s going on Anthropic?
Anthropic is partnering with @SpaceX to run models on their compute, exciting!
Claude Fable 5 from @AnthropicAI is live on OpenRouter! Anthropic's most capable coding model, built for long-running, ambiguous work: legacy migrations, gnarly production bugs and async sessions that run for hours or days. SOTA on nearly all tested benchmarks.
有讨论指出 Claude Fable 5 在用于前沿 LLM 开发时存在不易察觉的能力限制。该话题在 Hacker News 上仅获 3 分且暂无评论,正文未给出具体版本号、参数或基准数据。
模块化内核团队正在快速推进 M3 🚀 开放权重几天后发布——届时可直接在 @Modular 上运行。 很期待这个。
Our kernel team has been deep in MiniMax M3 all week. The 1M-token context and native multimodality make it a hard model to serve well, which is exactly the kind of problem we like! When the open weights drop in the next few days, you'll be able to run it on Modular right away. Stay tuned for @MiniMax_AI x Modular.
When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model’s capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic.
推荐理由:材料呈现模型在敏感领域被悄然降能的机制,可供观察厂商在能力与安全之间的取舍方式。



It's finally out!!! @METR_Evals found that more than half of SWEBench results is unmergeable slop. FrontierCode represents over 1000+ hours of maintainer validated software engineering work most frontier models cannot yet solve, much less solve with high quality. Cog had IOI Gold medalists and top code maintainers Look At The Data — FrontierCode includes 3000+ rubrics covering code quality and anticheat reward hacking plaguing other benchmarks. FC Diamond is so hard that Opus 4.8 scores 13.8%. Three eras of AI coding : Three eras of benchmarks 2021 • Autocomplete : HumanEval 2023 • Passing Tests: SWEBench, TerminalBench 2026 • Maintainable Code: FrontierCode to me the most beautiful chart when I requested a special historical run into all extant old models, the data was finding that the easiest third of FC tasks (in FC Extended) were rapidlly and suddenly solved over late 2025 - Opus almost doubled from a 41% pass rate to 74% in 4 months. This describes the "WTF happened in Dec 2025" vibe shift that a lot of folks from @dhh to @karpathy have called out: it is the difference between getting 95% success in 2 rerolls vs 6, making it finally feasible to go up the next layer of abstraction in agentic coding, eg @GeoffreyHuntley's ralph loops or @bcherny's /goals or @steipete's "loops that prompt your agents" without fearing too much that things go off the rails. My guess: as AI accelerates from here, each FrontierCode tier will saturate in sequence, hopefully ~annually. I've already asked the team to prepare FrontierCode 2027.... The old mountains will be destroyed. Their rubble becomes regolith. And from that regolith, the next model forest grows. Circle of life.
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time. I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!
推荐理由:转发的评测者认为官方榜单未体现 Fable 5 的进步幅度,可作为判断这次模型跃迁的定性参考。
为与 @Supplyaiusa 🤝 @HKGoodFortune 的合作感到自豪。
Maison Solutions Inc. (Nasdaq: MSS) has announced a strategic collaboration with @Supplyaiusa and @MiniMax_AI to explore AI-native food supply chain solutions. Together, the parties aim to bring AI closer to real food retail and supply chain operations. Read more: accessnewswire.com/newsroom/… $MSS #MaisonSolutions #SupplyAi #MiniMax #FoodSupplyChain #ArtificialIntelligence #RetailTech #SupplyChainInnovation
Big step for SupplyAi. We’re excited to be part of the strategic collaboration announced by @HKGoodFortune (Nasdaq: MSS) with @MiniMax_AI to explore AI-native food supply chain solutions. We’re building toward a future where business data, AI agents, and physical execution are more deeply connected across the food supply chain. Read more: accessnewswire.com/newsroom/… #SupplyAi #MaisonSolutions #MiniMax #FoodSupplyChain #AINative #AgenticAI #SupplyChainTechnology #RetailTech
👀 @MiniMax_AI
Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.
推荐理由:Karpathy 以第一手使用体验判断 Claude Fable 5 属版本级跃迁,并点出发布期安全阈值偏严的取舍。