Anthropic 发布 Claude Fable 5 与 Claude Mythos 5
Anthropic 发布 Mythos 级模型 Claude Fable 5,并同期向少数网络安全与基础设施合作方开放解除部分安全限制的 Claude Mythos 5。
推荐理由:原文给出两款新模型的能力定位、分级开放方式与安全回退机制,可对照理解前沿模型的分层发布思路。
AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。
当前仅显示精选新闻Anthropic 发布 Mythos 级模型 Claude Fable 5,并同期向少数网络安全与基础设施合作方开放解除部分安全限制的 Claude Mythos 5。
推荐理由:原文给出两款新模型的能力定位、分级开放方式与安全回退机制,可对照理解前沿模型的分层发布思路。
Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-recording-to-code, and more 🔹Plugins for stocks, financial reports, academic papers, with more coming 🔹Supports the ACP protocol, and works with JetBrains, Zed, and more 🔹Hooks for custom tools and workflows Try it with Kimi K2.6 👉 kimi.com/code Issues, plugin ideas, and PRs welcome! Community feedback helps shape what ships next.🚀
推荐理由:原文列出升级后的零配置安装、视频上下文与插件能力,读者可据此判断编码智能体的门槛变化。
OpenAI 高管提出「Chat 已死」,公司正推进 ChatGPT 诞生以来最大规模改版,目标是从聊天机器人变成个人 Agent 式的超级应用。改版由 Codex 承担 Agent 能力,其周活已超 500 万、非开发者用户占 20%,并新增可直接操作电脑、并行运行多个 Agent 等能力。动因是 ChatGPT 虽在 5 月突破 10 亿月活但多数用户免费,企业端支出正被 Claude 抢占。
推荐理由:文章梳理了 OpenAI 把 ChatGPT 从聊天框转向 Agent 超级应用的路线,并用 Codex 与 Claude 的数据呈现其转型压力。
Anthropic 研究院发布长文《当 AI 开始构建自己》,用公开基准和此前未披露的内部数据说明 AI 已在加速 AI 系统自身的开发。文中称 Anthropic 工程师平均每季度交付的代码量是 2021 至 2025 年间的 8 倍,Claude 在最开放任务上的成功率在 2026 年 5 月达到 76%,六个月内提高 50 个百分点。
推荐理由:Anthropic 用内部数据展示 Claude 在写代码和做研究上的进展,读者可据此理解递归自我改进这一趋势的早期证据。
Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:文中并列了编码速度、任务时长与代码占比等具体数字,可用来观察 AI 自主编码能力的演进节奏。
Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:汇总了 Anthropic 博客关于递归自我改进的关键数据与三种未来路径,可据此判断自动化编码的推进速度。




Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…
推荐理由:转述 Anthropic 内部数据,读者可据此了解递归自我改进讨论背后的具体加速指标。
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days
推荐理由:官方给出三项前沿能力与基准成绩,并说明权重和 MSA 算子的开源时间,读者可据此判断开放节奏与可用范围。
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days
推荐理由:MiniMax M3 在 Arena 前端编码榜位列第 7,与 GLM-5.1 接近,读者可据此比较同级价位模型的能力与定价。
Hugging Face 将官方命令行入口 hf CLI 重构为同时服务人类与编码智能体的工具,agent 模式下自动输出 TSV、不截断数据,并附带可直接执行的下一步命令提示。
推荐理由:官方给出 hf CLI 智能体模式的设计与基准数据,读者可据此了解编码 agent 调用 Hub 时的 token 开销差异。
Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days
推荐理由:引用的发布信息列出 M3 在编码、智能体与 1M 上下文上的基准成绩,可供判断这款开源权重模型的能力组合。
推荐理由:多数团队会堆叠路由器、微调与多模型,OpenAI 的数据智能体却只用单模型加 13 个工具跑通 90,000 张表,形成可对比的工程取舍。
Introducing Pareto Code: a new, free, experimental coding router Set `min_coding_score` in your request and route to the cheapest code-capable model that clears your bar, ranked by @ArtificialAnlys. See the Pareto frontier shifting in real time👇
推荐理由:请求可按代码能力门槛自动筛选最便宜可用模型,为控制编码模型成本提供了一种可配置的路由思路。
微软 AI 在 Build 上发布七个全新 MAI 模型,官方称并非简单迭代,而是从零开始、干净数据血统、零蒸馏训练的一整个家族,涵盖推理、编码、图像、转录与语音并各有 Flash 版本。
Seven new models launching at Build: let’s go! Reasoning. Code. Image. Transcribe. Voice. Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models Thread 🧵 #MSBuild
推荐理由:文中梳理了七个 MAI 模型的任务分工与基准数字,可据此了解微软从零训练、任务专精的模型家族路线。
Anthropic Claude Code 团队负责人复盘了智能体编码成为默认工作方式后对工程流程的重写:规划从六个月路线图改成 just-in-time,遇到问题先问 Claude 并追问能否自动化,代码评审由 Claude 处理风格、缺陷和测试,人只在法律风险、信任边界与安全代码、产品判断等需要领域专长处介入。
推荐理由:Claude Code 团队把规划、上下文获取、评审与分工按智能体编码重写,并给出可对照的前后变化与跟踪指标。
推荐理由:原文给出可一次安装的角色化插件、接入应用与技能数量,读者可据此判断 Codex 能覆盖哪些岗位工作。
We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install openai-codex Go build with it!!
推荐理由:SDK 把 Codex 从浏览器里的 IDE 变成可嵌进脚本的可编程基础设施,并复用现有认证。
Codex now has more than 5M weekly active users. But the bigger story is what people are using it for: not just writing code, but getting more work done across research, analysis, content, and operations. Our new report on how Codex is becoming a productivity tool for knowledge work: openai.com/index/codex-for-k…
推荐理由:OpenAI 公布 Codex 用户用途数据,知识工作者采用速度已超过开发者,可观察 AI 生产力工具的实际使用迁移。
We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install openai-codex Go build with it!!
推荐理由:原文给出 Codex Python SDK 的安装命令与复用登录态方式,读者可据此把它嵌入 Python 应用和工作流。
OpenAI 自家报告显示,Codex 周活跃用户已突破 400 万,较 2 月增长 5 倍。知识工作者目前占其中五分之一,增速为开发者的 3 倍。该报告由 Axios 首发。
推荐理由:OpenAI 报告显示 Codex 的知识工作者用户增速快于开发者,可作观察编码工具向非开发者扩散的参照。