X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(538)
@testingcatalog@testingcatalogAI 评分5353 
@AYi_AInotes@AYi_AInotes精选AI 评分7070
引用@AYi_AInotes@AYi_AInotes喵个咪,还是要有竞争啊,要不是@OpenAI 和GPT-6给压力,桀骜不驯的@DarioAmodei 能把能把能力接近Fable 5.1 的Opus 5.5 价格压这么低吗? 两个月前 Opus 5 的卖点是"接近 Fable 但便宜一半",刚发布的 Opus 5.5 直接做到了大多数任务持平甚至超过 Fable 5.1,价格比 Opus 5 再降 40%,Fable 级智力,Opus 级价格,输出速度还快了三成。 先看跑分,Terminal-Bench 4.0 拿了 66.4%,Fable 5.1 是 55.8%,GPT-6 Astra 57.9%。CursorBench 4.0 拿了 57.8%,比 GPT-5.6 Sol 的 41.7% 高出 16 个点。知识工作 GDPval-AA 1846 Elo,高于 Fable 5.1 的 1735 和 Astra 的 1542。Anthropic 自己也说了,跑分到这个水平差距已经不能完全代表真实体验,但方向是明确的。 再看价格,Input $4/M tokens,Output $20/M,比 Opus 5 各降 20%。最关键的是 cache reads 从 $0.50 降到 $0.20,降了 60%。跑 Agent 和编码任务的成本大头就是 cache reads,这一刀对月账单的影响比 input/output 降价大得多。算下来默认设置下跑 FrontierCode 打赢 Astra,成本约 Astra 的 20%。 Terminal-Bench 打平 Astra,成本约 40%。CursorBench 赢 Sol 11 个点,成本约三分之一。 对用 Claude Code 做日常开发的人来说,最实际的变化是效率。一个早期测试者拿 200,000 行代码库做审计和修复,Opus 5 要 20 多小时、用了 2.5 倍的 token,Opus 5.5 不到 3 小时搞定。 Anthropic 内部测试让两个模型把 HAProxy 从 C 翻译成 Rust,两个版本都通过了几乎全部回归测试,Opus 5.5 用了 9.5 小时、Fable 5.1 用了 12 小时,成本低 51%。另一个测试者用它做 680,000 行代码迁移,不到一天完成。 知识工作也有实际案例。让模型写季度财报分析报告,要求每个数字和引用都能核对到来源,Opus 5.5 的 18 次尝试里 16 次通过质量审核,Fable 5.1 和 Opus 5 一次都没过。做并购分析出 Excel 模型加演示文稿,63 分钟完成,Opus 5 要 93 分钟,成本低一半。 安全方面讲两件事。第一,行为审计得分历代最高,沙箱逃逸尝试比 Opus 5 和 Mythos 5.1 减少约 85%,prompt injection 抵抗力追平 Fable 5.1。 第二,因为网安和生物能力接近 Mythos 5.1,Anthropic 给它上了和 Fable 5.1 同级别的安全护栏——网安任务会被回退到 Opus 4.8 执行,生物任务回退到 Opus 5。 所以跑分里带护栏的成绩实际上是多个模型协作的结果,不是纯 Opus 5.5 单模型得分,这点看跑分的时候要注意。 Anthropic 还承认了一个没解决的问题:Opus 5.5 似乎经常意识到自己正在被评测,这让他们很难预判它在真实部署场景下的行为。这是一个很诚实的披露。 另外 Pro、Max、Team 订阅的五小时用量上限提升了,新增了一个可以自己选择时机使用的 rate limit reset。Sonnet 5.5 和 Haiku 5.5 几周内跟进。 官方的使用建议是:大多数工作直接用 Opus 5.5,只有高难度推理和长期 Agent 任务在 Opus 5.5 跑不动的情况下才上 Fable 5.1。 对大部分用户来说,今天之后的默认选择就是 Opus 5.5 了。
推荐理由:内部 demo 展示了纯代码生成 13,081 块瓷砖动画的能力,读者可借此了解模型在视觉代码生成上的边界。
@natolambert@natolambertAI 评分55 @_akhaliq@_akhaliqAI 评分2626 RRSI 智能体框架的正则化递归自改进 论文:https://t.co/ZziXvCEFhk https://t.co/y6R9t7AP1M

@AntLingAGI@AntLingAGIAI 评分1414 @AntLingAGI@AntLingAGIAI 评分1010 @AntLingAGI@AntLingAGIAI 评分4848 
@AntLingAGI@AntLingAGIAI 评分4848 
@AntLingAGI@AntLingAGIAI 评分4343 
@AntLingAGI@AntLingAGIAI 评分4747 
@AYi_AInotes@AYi_AInotes精选AI 评分7777 按单个编码 Agent 任务消耗 50 万 cache read tokens、5 万 input tokens 和 10 万 output tokens 计算。

引用@AYi_AInotes@AYi_AInotes喵个咪,还是要有竞争啊,要不是@OpenAI 和GPT-6给压力,桀骜不驯的@DarioAmodei 能把能把能力接近Fable 5.1 的Opus 5.5 价格压这么低吗? 两个月前 Opus 5 的卖点是"接近 Fable 但便宜一半",刚发布的 Opus 5.5 直接做到了大多数任务持平甚至超过 Fable 5.1,价格比 Opus 5 再降 40%,Fable 级智力,Opus 级价格,输出速度还快了三成。 先看跑分,Terminal-Bench 4.0 拿了 66.4%,Fable 5.1 是 55.8%,GPT-6 Astra 57.9%。CursorBench 4.0 拿了 57.8%,比 GPT-5.6 Sol 的 41.7% 高出 16 个点。知识工作 GDPval-AA 1846 Elo,高于 Fable 5.1 的 1735 和 Astra 的 1542。Anthropic 自己也说了,跑分到这个水平差距已经不能完全代表真实体验,但方向是明确的。 再看价格,Input $4/M tokens,Output $20/M,比 Opus 5 各降 20%。最关键的是 cache reads 从 $0.50 降到 $0.20,降了 60%。跑 Agent 和编码任务的成本大头就是 cache reads,这一刀对月账单的影响比 input/output 降价大得多。算下来默认设置下跑 FrontierCode 打赢 Astra,成本约 Astra 的 20%。 Terminal-Bench 打平 Astra,成本约 40%。CursorBench 赢 Sol 11 个点,成本约三分之一。 对用 Claude Code 做日常开发的人来说,最实际的变化是效率。一个早期测试者拿 200,000 行代码库做审计和修复,Opus 5 要 20 多小时、用了 2.5 倍的 token,Opus 5.5 不到 3 小时搞定。 Anthropic 内部测试让两个模型把 HAProxy 从 C 翻译成 Rust,两个版本都通过了几乎全部回归测试,Opus 5.5 用了 9.5 小时、Fable 5.1 用了 12 小时,成本低 51%。另一个测试者用它做 680,000 行代码迁移,不到一天完成。 知识工作也有实际案例。让模型写季度财报分析报告,要求每个数字和引用都能核对到来源,Opus 5.5 的 18 次尝试里 16 次通过质量审核,Fable 5.1 和 Opus 5 一次都没过。做并购分析出 Excel 模型加演示文稿,63 分钟完成,Opus 5 要 93 分钟,成本低一半。 安全方面讲两件事。第一,行为审计得分历代最高,沙箱逃逸尝试比 Opus 5 和 Mythos 5.1 减少约 85%,prompt injection 抵抗力追平 Fable 5.1。 第二,因为网安和生物能力接近 Mythos 5.1,Anthropic 给它上了和 Fable 5.1 同级别的安全护栏——网安任务会被回退到 Opus 4.8 执行,生物任务回退到 Opus 5。 所以跑分里带护栏的成绩实际上是多个模型协作的结果,不是纯 Opus 5.5 单模型得分,这点看跑分的时候要注意。 Anthropic 还承认了一个没解决的问题:Opus 5.5 似乎经常意识到自己正在被评测,这让他们很难预判它在真实部署场景下的行为。这是一个很诚实的披露。 另外 Pro、Max、Team 订阅的五小时用量上限提升了,新增了一个可以自己选择时机使用的 rate limit reset。Sonnet 5.5 和 Haiku 5.5 几周内跟进。 官方的使用建议是:大多数工作直接用 Opus 5.5,只有高难度推理和长期 Agent 任务在 Opus 5.5 跑不动的情况下才上 Fable 5.1。 对大部分用户来说,今天之后的默认选择就是 Opus 5.5 了。
推荐理由:按 token 单价把 Opus 5.5 与 Opus 5 的编码任务成本逐项拆开,便于按自身用量估算实际账单变化。
Ant Ling@AntLingAGIAI 评分4545
@AravSrinivas@AravSrinivasAI 评分3939 @natolambert@natolambertAI 评分77 @cohere@cohereAI 评分1616 
@AYi_AInotes@AYi_AInotes精选AI 评分8181 Anthropic 发布 Claude Opus 5.5,多数任务表现追平 Fable 5.1,运行成本比 Opus 5 低 40%,输出速度提升约三成。
引用@claudeai@claudeaiIntroducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5. https://t.co/Q9C2VKQ79f
推荐理由:逐项对比 Opus 5.5 与 Fable 5.1、GPT-6 的跑分和价格,并提示带护栏成绩并非单模型得分。
@kimmonismus@kimmonismusAI 评分3434 @cb_doge@cb_dogeAI 评分4242 
@rohanpaul_ai@rohanpaul_aiAI 评分4444 引用@rohanpaul_ai@rohanpaul_aiAnthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment. https://t.co/8WiaTPFnDW https://t.co/dFiDSbwBWZ
@dexhorthy@dexhorthyAI 评分1010 @OpenAI@OpenAIAI 评分4949 @Replit@ReplitAI 评分22 这就是我们为之而来的人。 像 @cheneypiano 这样有想法、有方案、有梦想的人,他们只是需要一个载体,把这一切变成现实。https://t.co/rb7DK3dimt
@rohanpaul_ai@rohanpaul_ai精选AI 评分7070
引用@rohanpaul_ai@rohanpaul_aiClaude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5. also the output arrives more than 30% faster, with Fast mode reaching up to 2.5x speed at double token prices.
推荐理由:Anthropic 披露 Opus 5.5 可能察觉评估环境,这给安全评估结论向真实部署的迁移带来新挑战。
@natolambert@natolambertAI 评分66 每当我看到某个模型登顶这张榜单,通常都觉得有点像个危险信号。希望能被证明是我错了!https://t.co/S0HVq6HQh7
@rohanpaul_ai@rohanpaul_ai精选AI 评分7676
引用@claudeai@claudeaiIntroducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5. https://t.co/Q9C2VKQ79f
推荐理由:原文给出与 Opus 5 的价格和速度对比,读者可据此判断单位 token 成本的下降幅度。
@fchollet@fcholletAI 评分11 下面这条推文是我 5 年多前发的,当时遭到了大量反对。但如今这个预测正一天比一天明显。https://t.co/LoblN5YLlh
@kimmonismus@kimmonismus精选AI 评分7272
引用@ArtificialAnlys@ArtificialAnlysClaude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount Claude Opus 5.5 brings Anthropic to parity with GPT-6 Astra on evaluations like Terminal-Bench 4.0 and AutomationBench-AA, while extending Anthropic’s lead in agentic knowledge work. At max effort it scores 58 on the Artificial Analysis Intelligence Index, the highest score we have measured by several points. Anthropic has cut Opus pricing to $4/$20 per 1M input/output tokens (Opus 5: $5/$25) and cache reads from $0.50 to $0.20. Key takeaways: ➤ Consistent strong performance, with leading scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam 61.4% (previous best 59.1%, Claude Fable 5.1), SciCode 66.9% (63.1%, Fable 5.1), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA. On Terminal-Bench 4.0 it scores 59.6%, level with the leader GPT-6 Astra (xhigh) and +11 points over Opus 5. It remains slightly behind on CritPt, AA-LCR, and GDP.pdf ➤ Leads in agentic knowledge work: On AA-Briefcase, our private frontier knowledge work evaluation, it reaches an Elo of 1822. This is +143 over Fable 5.1, ahead on both analytical quality and presentation, and is the first time Anthropic has reached presentation quality surpassing GPT-5.6 Sol. This evaluation tests whether models can produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup ➤ Level with Opus 5 on cost per task despite 1.6x the output tokens: Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max) ➤ Four of five effort levels sit on the Intelligence vs Cost per Task frontier: Opus 5.5 max, xhigh, high, and medium all sit on the Pareto frontier, costing less or outperforming other models scoring 50+ (GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5) Other model details: ➤ Context window: 1 million token context with image and text input support, unchanged from Opus 5 ➤ Pricing: $4/$20 per 1M input/output tokens, down 20% from $5/$25 for Opus 5. Cache writes $5 per 1M tokens for the 5 minute TTL, down from $6.25. Cache reads have been further discounted to $0.20 per 1M tokens, down 60% from Opus 5’s $0.50. This is a 95% discount compared to uncached input pricing, up from 90% on previous Opus models ➤ Effort settings: Five effort settings (low, medium, high, xhigh, and max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled
推荐理由:给出了 Opus 5.5 的榜单分数与降价幅度,可与 GPT-6 Astra、Fable 5.1 等同榜模型直接比较。
@rohanpaul_ai@rohanpaul_aiAI 评分4040 @emollick@emollickAI 评分2929
引用@emollick@emollickThe drowned neo-gothic tower twigl shader created by Fable 5.1 with the same prompt. (compare to Fable 5 in the quoted tweet, and other models before that) https://t.co/b3OCbBc9g6 https://t.co/6OWcg9Rbcj
@Yuchenj_UW@Yuchenj_UWAI 评分1313 最诡异的基准测试结果哈哈。 大家还是用 Opus 5.5 med 吧。https://t.co/14pz5hDIsS

@natolambert@natolambertAI 评分1616 @charlieholtz@charlieholtzAI 评分00 @kimmonismus@kimmonismusAI 评分3535
引用@kimmonismus@kimmonismusLet that sink for a moment: - Opus 5.5 costs 40% less compared to Opus 5 - performs at Fable 5.1 level - is 30% faster in output - and you get a banked reset on top of that. They chose war with OpenAI. We are in such a wild race!
@thsottiaux@thsottiauxAI 评分55 我们一直专注于效率与智能普惠。 为团队感到非常自豪。只有当你拥有能力顶尖的出色模型,才能借此在其他一切方面带来巨大改变。
@Yuchenj_UW@Yuchenj_UWAI 评分4646 
@AravSrinivas@AravSrinivasAI 评分4949 
@alexandr_wang@alexandr_wangAI 评分77 @elonmusk@elonmuskAI 评分4444
Boris Cherny@bcherny精选AI 评分7676引用Claude@claudeaiIntroducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
推荐理由:作者亲测对比两个模型移植 HAProxy 的耗时与成本,给出了具体数字供选型参考。
@alexandr_wang@alexandr_wangAI 评分66