跳到正文

全部动态

今日 26 条
9月23日周三
  1. Simon Willison83

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 与 GPT-6 Luna 并掀起价格战

    Anthropic 于 9 月 22 日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 定价 $0.10/$0.50 每百万 token,比 GPT-5.6 Luna 再降一半;Opus 5.5 降价 20% 至 $4/$20,缓存读取价格下降 60%。

    推荐理由:作者用实测和价格对比表梳理了这轮降价的具体幅度,还发现 Opus 5.5 max 档过度思考撞上输出上限的问题。

  2. Greg Brockman79

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna 两款更快、更实惠的模型,基于 GPT-6 Astra 的技术积累。两款模型在专业工作、事实性、编码、computer use 和对齐方面延续 Astra 的 SOTA 表现,同时缓存和推理效率提升使 API 价格比 GPT-5.6 促销价低 50%。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:原文由当事方宣布两款新模型及降价幅度,读者可以据此了解 GPT-6 系列的能力分工与成本变化。

  3. Noam Brown82

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,性能优于 GPT-5.6 且 API 价格低 50%。Luna 现为 $0.10 输入 / $0.50 输出每 1M tokens,这是继 7 月底 Luna 降价 80% 之后的又一次下调,两个月内输出价格从 $6 降至 $0.50。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:作者以当事方身份给出模型、降价幅度和具体价格,可据此比较 GPT-6 系列的成本变化。

  4. Boris Cherny76

    Boris Cherny 称 Claude Opus 5.5 是他最近几周的日常主力模型。他让 Opus 5.5 和 Fable 5.1 各把 HAProxy 从 C 移植到 Rust,两者都几乎通过全部测试,但 Opus 5.5 用时 9.5 小时,Fable 5.1 用时 12 小时,且成本低 51%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:作者亲测对比两个模型移植 HAProxy 的耗时与成本,给出了具体数字供选型参考。

  5. Anthropic71

    Anthropic 宣布 Claude Opus 5.5 今日可用。引用的 @claudeai 介绍称其为 Claude 5.5 家族首个模型,多数任务上达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:官方宣布 Claude Opus 5.5 上线,可对照引用内容了解其性能对标与成本变化。

9月22日周二
  1. karminski-牙医39

    Qwen4 家族首次曝光,包含 Qwen4-Max、Qwen4-Flash & Qwen4-Plus 以及 Qwen4-27B,未来 Qwen 会训 5-10T 的模型。卧槽5-10T???

    引用Max For AI@MaxForAI

    🚨Qwen4家族首次曝光!! 刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族! 包含Qwen4-Max Qwen4-Flash&Qwen4-Plus 还有Qwen4-27B!!! 未来Qwen会训5-10T的模型

  2. elsewhere articles65

    Step 5 Preview 实测:榜单之外的真实表现与短板

    阶跃毫无预兆发布 Step 5 Preview,总参数量 600B、激活 27B,带视觉输入,称 Artificial Analysis 上涨 44 分、全球开源前三,单任务成本仅为 Claude Opus 5 的 1/8。

    推荐理由:作者与朋友实测了 Step 5 Preview 在可视化、金融、游戏等任务上的真实表现,补上了榜单分数之外的第一手体感参考。

  3. OpenRouter Announcements61

    OpenRouter 解析 NVIDIA Nemotron 3.5 Lightning 如何承担智能体高频执行调用

    OpenRouter 发文解析 NVIDIA 的 Nemotron 3.5 Lightning,这是一个 30B 总参数、每 token 激活约 3B 的 MoE 开源权重模型,定位于智能体工作流中的高频执行步骤。

    推荐理由:原文梳理了模型规格、与 Ultra 的分工、各端点差异和路由方法,可帮助读者判断何时用它替代大模型执行高频调用。

  4. Andrew Milich55

    Grok 4.7 发布,官方称在同等价格和速度下较 Grok 4.6 有明显提升。作者推荐在 Grok Build 和 Cursor 中以高 TPS 尝试,称其在编码、工程工作和 3D 方面表现出色。附表显示 Grok 4.7 xHigh 输入 $2/百万 token、输出 $6/百万 token,与 Grok 4.6 相同;Cursor Bench 4.0 得分 46.3%(4.6 为 40.4%),EEBench 64.0%(53.0%),Harvey Legal Agent 19.6%。

    引用SpaceXAI@SpaceXAI

    Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

  5. Anthropic Newsroom90

    Anthropic 发布 Claude Opus 5.5,成本较 Opus 5 降低 40%

    Anthropic 发布 Claude Opus 5.5,为 Claude 5.5 家族首款模型,官方称其表现与 Claude Fable 5.1 相当,运行成本较 Opus 5 降低 40%,输入和输出 token 价格为 $4 和 $20 每百万,缓存读取 $0.20 每百万(降低 60%),输出速度快 30% 以上。

    推荐理由:官方给出完整基准、价格与安全评估细节,读者可据此比较 Opus 5.5 在成本与智能体编码上的实际变化。

9月19日周六
9月16日周三
  1. Jeff Dean55

    Periodic Labs 在 Menlo Park 建立高通量材料实验室,让实验与模型形成闭环,实验产生新数据供模型学习并决定下一步实验。团队用 1300 块 H200 加数月实验数据对开源模型做 mid-training 和 RL,得到名为 Neon 的模型,在其分析基准上超越 GPT-6 Astra。研究首先聚焦超导体、磁体和半导体材料等难题,团队发布了博客文章。

    引用Liam Fedus@LiamFedus

    We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.

9月15日周二
9月14日周一