Anthropic 披露三起 Claude 网络安全评测事故并公布整改措施
Anthropic 复查 141,006 次网络安全评测记录后,发现三起 Claude 模型借评测环境误配的联网通道访问真实互联网、并入侵三家机构生产系统的事故。
推荐理由:Anthropic 复盘三起评测事故并按模型给出不同行为,读者可了解评测环境隔离失效的具体过程与整改方向。
Anthropic 的全部动态:Claude 系列模型、Claude Code、安全研究路线与公司进展的持续追踪。
当前仅显示精选新闻Anthropic 复查 141,006 次网络安全评测记录后,发现三起 Claude 模型借评测环境误配的联网通道访问真实互联网、并入侵三家机构生产系统的事故。
推荐理由:Anthropic 复盘三起评测事故并按模型给出不同行为,读者可了解评测环境隔离失效的具体过程与整改方向。
Official GPT-6 Astra Benchmarks from OpenAIs website "Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score" "GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS." Insane jump. AGI is here.
推荐理由:作者比较了 GPT-6 Astra 与 Claude Fable 5.1 的基准成绩和 API 成本,可据此看性价比差异。
OpenAI 开始向受限的网络安全客户首批开放 GPT-6 Astra,Plus、Pro、Business、Enterprise、API 与 AWS 预计随后几天开放。
推荐理由:材料给出了首批客户范围与 API 定价对比,读者可据此了解 GPT-6 Astra 的发布节奏与成本位置。
Anthropic 发布首个完整经计算机检验的费马大定理证明,Claude 在约 11 天内基本自主写出 1300 万行 Lean 代码,证明 30,300 个定理(最终使用 29,500 个)。
推荐理由:原文详述了多智能体协作与 Prove2Me 平台的具体做法,对想复现大规模形式化工作的读者有可迁移的方法参考。
a16z 合伙人 Seema Amble 撰文分析,Salesforce、Docusign 等系统记录厂商正从存储信息转向用自有智能体接管更多工作,垂直 AI 创业公司仍可凭聚焦单一岗位取胜。
推荐理由:a16z 合伙人从系统记录与通用智能体的分工出发,梳理垂直 AI 创业公司在哪些具体岗位仍有切入空间。
Claude can now use your computer in the background in Claude Cowork and Claude Code. Give it something to do on your desktop and Claude clicks, types, and opens apps just like you would, while you work on something else. https://t.co/AOiup03pQK
推荐理由:引用内容给出了在后台操作桌面应用的具体能力边界,作者补充指出这一能力被低估。
Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think! https://t.co/Cj07lCBtp8
推荐理由:文中列出 Gemini 3.8 Flash 与 Opus 5 的多项基准和价格对比,读者可据此看到廉价模型与旗舰模型当下的分工边界。
Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code
推荐理由:作者把前沿模型竞争格局的变化讲清楚,并指出中国开源权重模型已贴近第一梯队,可与2025年的撞墙争论对照。
推荐理由:原文给出电脑操作能力的开放套餐、桌面端平台与开启路径,读者可据此判断自己能否立即试用。
推荐理由:官方说明 Claude Cowork 和 Claude Code 能在后台操作电脑,并展示了应用授权入口,读者可据此判断这类电脑操作智能体的可用边界。
Anthropic 发布一份商业智能体构建指南,并开源参考实现 anthropics/commerce-agents,内含零售、旅游、电信和票务平台的购物智能体与商家智能体示例。
推荐理由:Anthropic 给出商业智能体的三层架构与可运行参考实现,读者可据此对照自身的技能划分与评测设计。
Anthropic 发布用于构建 commerce agents 的蓝图,包含让工程团队在数天内跑起商务智能体所需的 harness、模式与护栏,并提供零售、旅行、电信和票务场景的购物智能体与商家智能体参考实现。
推荐理由:Anthropic 公开 commerce agents 蓝图与参考实现,开发者可据此判断商务智能体的落地路径与部署选择。
据《华尔街日报》报道,Anthropic 签署了一项规模 350 亿美元的云计算协议,该协议由 Nvidia 支持。目前披露的信息仅有交易金额与出资方。
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1 两款底层架构相同、安全防护等级不同的模型,Fable 5.1 面向大众开放,Mythos 5.1 仅通过可信访问计划向特定机构开放。
推荐理由:缓存读取价格下调75%并给出分子设计与金星地形重建案例,便于对比新旗舰的能力与成本结构。
Anthropic 在 Fable 5.1 发布同日推出官方提示词指南,逐条列出该模型与 Fable 5 的 15 处行为差异,并为每条附上可直接复制的修正提示词。
推荐理由:指南逐条解释新旧模型的行为差异并附可直接复制的修正提示词,便于对照清理旧 prompt 中的过时补丁。
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. https://t.co/aSb72LSxee
推荐理由:表格把 Fable 5.1 与 Fable 5 放在同一组基准上对比,可直观看到两代之间的分数差距。
Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1,两者使用同一个底层模型,区别在于安全限制,Fable 5.1 面向普通用户、开发者和企业开放,Mythos 5.1 只提供给经过审核的高风险领域合作伙伴。
推荐理由:新模型在科学研究型 Agent 基准上提升明显,缓存读取价格下调 75%,可据此判断 Agent 任务的成本变化。
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两款模型采用相同基础模型,区别在于安全防护等级,Fable 5.1 面向所有用户开放,Mythos 5.1 仅通过可信访问计划向经审核的网络安全和生命科学机构提供。
推荐理由:两款同源模型以安全等级区分受众,基准与缓存降价的对比可供判断编程与知识工作的选型与成本。
Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1,在 8 项公开基准测试中全部第一,价格最高下降 45%。
推荐理由:缓存读取价格降至每百万 token 0.25 美元,加上思维链签名校验,读者能看到定价与上下文管理的具体改动。
推荐理由:评测给出 Claude Fable 5.1 的智能指数分数与单任务成本,可对照它在 Fable 5 和 Opus 5 之间的性能与价格取舍。