跳到正文

Anthropic / Claude

Anthropic 的全部动态:Claude 系列模型、Claude Code、安全研究路线与公司进展的持续追踪。

当前仅显示精选新闻

最新精选

第 141–160 条 · 共 497 条
9月4日周五
  1. @kimmonismus75

    OpenAI 的 GPT-6 Astra 开始向有限机构推送,并将在未来几天面向所有 ChatGPT Plus、Pro、Business、Enterprise 用户以及 OpenAI API 和 AWS 开放。官方基准称 Astra 在 ARC-AGI-3 上取得 99.9%、在 ExploitBench 上取得 100%。作者表示该模型在各项基准上全面超过 Claude Fable 5.1,且成本更低。

    引用@kimmonismus@kimmonismus

    Official GPT-6 Astra Benchmarks from OpenAIs website "Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score" "GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS." Insane jump. AGI is here.

    推荐理由:作者比较了 GPT-6 Astra 与 Claude Fable 5.1 的基准成绩和 API 成本,可据此看性价比差异。

  2. Anthropic Research80

    Anthropic:Claude 用 11 天完成费马大定理首个完整机器验证证明

    Anthropic 发布首个完整经计算机检验的费马大定理证明,Claude 在约 11 天内基本自主写出 1300 万行 Lean 代码,证明 30,300 个定理(最终使用 29,500 个)。

    推荐理由:原文详述了多智能体协作与 Prove2Me 平台的具体做法,对想复现大规模形式化工作的读者有可迁移的方法参考。

9月3日周四
  1. @bcherny66

    Claude Cowork 和 Claude Code 新增后台电脑使用功能,Claude 可以像用户一样在桌面点击、输入并打开应用,同时用戶可以处理别的事情。Boris Cherny 转发该消息并评论称,后台电脑使用这一能力被低估。

    引用@claudeai@claudeai

    Claude can now use your computer in the background in Claude Cowork and Claude Code. Give it something to do on your desktop and Claude clicks, types, and opens apps just like you would, while you work on something else. https://t.co/AOiup03pQK

    推荐理由:引用内容给出了在后台操作桌面应用的具体能力边界,作者补充指出这一能力被低估。

  2. @AYi_AInotes65

    Google 发布 Gemini 3.8 Flash,官方称其在智能体和编码能力上继续提升,是 6 周内第三个更新的 Flash 模型。作者对比称其价格约为 Claude Opus 5 的 1.5 折,在法律 10.0% 对 6.7%、长视频 87.8% 对 75.4% 上反超;但 OSWorld 电脑操作 59.0% 对 75.4%、通用 Agent 规划 19.1% 对 51.8% 仍落后。

    引用@OfficialLoganK@OfficialLoganK

    Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think! https://t.co/Cj07lCBtp8

    推荐理由:文中列出 Gemini 3.8 Flash 与 Opus 5 的多项基准和价格对比,读者可据此看到廉价模型与旗舰模型当下的分工边界。

  3. @kimmonismus67

    前沿模型竞争格局在很短时间内从 OpenAI 与 Anthropic 双强之争,扩展为 OpenAI、Anthropic、xAI、Meta 多方并跑、Google 重新加入的局面,中国开源权重模型也紧随其后。

    引用@ArtificialAnlys@ArtificialAnlys

    Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code

    推荐理由:作者把前沿模型竞争格局的变化讲清楚,并指出中国开源权重模型已贴近第一梯队,可与2025年的撞墙争论对照。

  4. @claudeai77

    Claude 现在可在 Claude Cowork 和 Claude Code 中后台使用电脑,用户给桌面任务后,Claude 会像用户一样点击、输入并打开应用,同时用户可以处理其他事情。官方展示的授权界面显示,电脑使用需要用户批准,允许 Claude 在后台或完全控制屏幕的情况下查看和操作已批准的应用。

    推荐理由:官方说明 Claude Cowork 和 Claude Code 能在后台操作电脑,并展示了应用授权入口,读者可据此判断这类电脑操作智能体的可用边界。

  5. Claude Blog63

    Anthropic 发布 Claude 商务智能体蓝图,含购物与商家智能体参考实现

    Anthropic 发布用于构建 commerce agents 的蓝图,包含让工程团队在数天内跑起商务智能体所需的 harness、模式与护栏,并提供零售、旅行、电信和票务场景的购物智能体与商家智能体参考实现。

    推荐理由:Anthropic 公开 commerce agents 蓝图与参考实现,开发者可据此判断商务智能体的落地路径与部署选择。

9月2日周三
  1. AI寒武纪 · 微信公众号77

    Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1 双旗舰模型,缓存读取价格降 75%

    Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1 两款底层架构相同、安全防护等级不同的模型,Fable 5.1 面向大众开放,Mythos 5.1 仅通过可信访问计划向特定机构开放。

    推荐理由:缓存读取价格下调75%并给出分子设计与金星地形重建案例,便于对比新旗舰的能力与成本结构。

  2. @AISafetyMemes77

    Fable 5.1 在 Terminal-Bench-Science 0.1 上得分 52.6%,超过 Fable 5 的 24.7%;在 Terminal-Bench 4.0 上以 55.8% 对 42.0% 领先,GDPval-AA v2 上为 1853 对 1723。引用内容称该模型在多项基准上刷新标准。作者指出这两代之间只相隔 2.5 个月,而不是数年。

    引用@claudeai@claudeai

    Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. https://t.co/aSb72LSxee

    推荐理由:表格把 Fable 5.1 与 Fable 5 放在同一组基准上对比,可直观看到两代之间的分数差距。

  3. 机器之心 · 微信公众号84

    Anthropic 发布 Claude Fable 5.1 与 Mythos 5.1,缓存读取降价 75%

    Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1,两者使用同一个底层模型,区别在于安全限制,Fable 5.1 面向普通用户、开发者和企业开放,Mythos 5.1 只提供给经过审核的高风险领域合作伙伴。

    推荐理由:新模型在科学研究型 Agent 基准上提升明显,缓存读取价格下调 75%,可据此判断 Agent 任务的成本变化。

  4. IT Home76

    Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1:性能超越前代,缓存读取费用下调 75%

    Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两款模型采用相同基础模型,区别在于安全防护等级,Fable 5.1 面向所有用户开放,Mythos 5.1 仅通过可信访问计划向经审核的网络安全和生命科学机构提供。

    推荐理由:两款同源模型以安全等级区分受众,基准与缓存降价的对比可供判断编程与知识工作的选型与成本。