跳到正文

#推理

今日 19 条
9月23日周三
  1. Noam Brown82

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,性能优于 GPT-5.6 且 API 价格低 50%。Luna 现为 $0.10 输入 / $0.50 输出每 1M tokens,这是继 7 月底 Luna 降价 80% 之后的又一次下调,两个月内输出价格从 $6 降至 $0.50。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:作者以当事方身份给出模型、降价幅度和具体价格,可据此比较 GPT-6 系列的成本变化。

9月22日周二
  1. Jeff Dean39

    感谢精彩的讨论,@dawnsongtweets!

    引用Dawn Song@dawnsongtweets

    I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8

  2. Andrew Milich55

    Grok 4.7 发布,官方称在同等价格和速度下较 Grok 4.6 有明显提升。作者推荐在 Grok Build 和 Cursor 中以高 TPS 尝试,称其在编码、工程工作和 3D 方面表现出色。附表显示 Grok 4.7 xHigh 输入 $2/百万 token、输出 $6/百万 token,与 Grok 4.6 相同;Cursor Bench 4.0 得分 46.3%(4.6 为 40.4%),EEBench 64.0%(53.0%),Harvey Legal Agent 19.6%。

    引用SpaceXAI@SpaceXAI

    Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

  3. Anthropic Newsroom90

    Anthropic 发布 Claude Opus 5.5,成本较 Opus 5 降低 40%

    Anthropic 发布 Claude Opus 5.5,为 Claude 5.5 家族首款模型,官方称其表现与 Claude Fable 5.1 相当,运行成本较 Opus 5 降低 40%,输入和输出 token 价格为 $4 和 $20 每百万,缓存读取 $0.20 每百万(降低 60%),输出速度快 30% 以上。

    推荐理由:官方给出完整基准、价格与安全评估细节,读者可据此比较 Opus 5.5 在成本与智能体编码上的实际变化。

9月21日周一
  1. OpenRouter Announcements68

    TypeSafe 决策模型 Jev 1.13 上线 OpenRouter,面向开发者解析用法

    TypeSafe 于 2026 年 9 月 15 日发布早期访问版决策模型 Jev(当前版本 1.13),现已可通过 OpenRouter 调用。Jev 是非生成式决策模型,只接受文本输入,返回 Choice、Score、Noul 三种类型化答案并附带校准概率,无自由文本输出。

    推荐理由:原文给出 Jev 三个返回原语、真实 API 响应和定价细节,开发者可据此判断是否用它替换现有 LLM 加正则的分类流程。

9月20日周日
  1. elsewhere articles65

    Sayash Kapoor 与 Arvind Narayanan 分析 AI 会不会让论文更快却让科学进步更慢

    Sayash Kapoor 和 Arvind Narayanan 在 2025 年发表的文章提出生产—进步悖论:全球论文数量约每 12 年翻一番,1900 至 2015 年间增长约 500 倍,但诺奖成果诞生于获奖前 20 年内的比例从 1970 年约 90% 降至 2015 年约 50%,科学进步相对投入明显放缓。

    推荐理由:文章把 AI 加速科研的讨论从模型能力转向注意力、激励、可复现性和人类理解等制度瓶颈,提供了评估 AI 科研工具的三个问题。

9月19日周六
9月18日周五
  1. Noam Brown51

    OpenAI 的 Noam Brown 在 Dwarkesh 播客中深谈多智能体、Navier-Stokes 与当前数学进展对自动化 AI 研究和递归自我改进的启示。讨论还涵盖如何在启动 RSI 前判断模型是否真正对齐,以及思维链退化、内外部模型差距等话题,并感谢 OpenAI 团队在多智能体方面的工作。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?

9月17日周四
9月16日周三
  1. Google Research43

    Google Research 提出 Retrieve-for-Train:用 RL 编译扩散模型绕过推理瓶颈

    Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现单次非自回归的查询扇出,无需推理时的 CoT 思考 token。该方法基于 Gemma3-4B 和 Qwen3-4B 微调,旨在解决零样本 LLM 在数据库感知查询分解中的复述坍缩与自回归延迟瓶颈。

9月15日周二
9月14日周一
9月12日周六
  1. Dwarkesh Patel56

    Dwarkesh 对谈 John Schulman、Beren Millidge 与 Charlie O'Neill:递归自我改进还有多远

    Dwarkesh Patel 与 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman、Baseten 模型训练负责人 Charlie O'Neill 长篇对谈,逐段讨论递归自我改进(RSI)最可能失败的技术原因、中国实验室的追赶路径、自动化 AI 研究者的训练方式以及长时程 RL 能否带来 AGI。

9月10日周四
9月9日周三
  1. Eric70

    OpenAI 宣布给出纳维-斯托克斯千禧年大奖难题的一个解,证明由一组智能体使用比 GPT-6 Astra 能力更强的 OpenAI 下一代模型产出。该问题关注纳维-斯托克斯方程描述的光滑三维流体运动是否会崩溃,约 90 年来未获解决。转发作者以在 OpenAI 工作的口吻感叹又是一个疯狂的日子。

    引用OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 官方宣布用下一代模型的智能体群产出纳维-斯托克斯千年问题证明,读者可以关注智能体做数学研究的这一路径。

9月3日周四
9月2日周三
  1. Hugging Face Blog52

    Allen AI 发布 BenchMIRT:在题目层面审计 LLM 基准实际测量的能力

    Allen AI 发布 BenchMIRT,一种基于多维 IRT 的方法,用于在单个题目层面审计 LLM 基准实际测量的能力。方法基于 100 个 LLM 在 16 个基准、超 34K 题目上的结果训练,未事先告知标签即独立恢复出安全与通用推理两个维度;分析发现 BBQ 和 WMDP 与通用推理的关联强于安全,HarmBench 的版权题也更接近推理。

8月26日周三
  1. Z.ai72

    智谱(Z.ai)发布 GLM-5.3-Flash,称具备有竞争力的价格与原生多模态能力,上下文窗口为 1M token,为 320B-A18B 模型并以 MIT License 开源权重。该模型此前曾以 Ox Alpha 名义预览,完全运行于中国 AI 芯片;现已在官方平台提供权重、API、Coding Plan、ZCode、Chat 和 AutoClaw 入口。

    推荐理由:官方公告同时给出价格定位、开源权重和芯片适配信息,读者可以据此评估它在现有工作流中的替换可能。

8月25日周二
  1. Hugging Face Blog63

    IBM 发布 Granite 4.2 推理模型家族并详解构建过程

    IBM 发布 Granite 4.2 密集 decoder-only 推理模型家族,含 3B、8B、30B 三个规格,基于 Granite-4.1 基座(约 15T tokens 预训练,上下文窗口扩至 512K),经 SFT 与多阶段 GRPO 强化学习训练,全部以 Apache 2.0 许可开源。

    推荐理由:IBM 官方详解 Granite 4.2 训练全程,从五阶段预训练到多阶段 RL 课程,可复用的训练细节较完整。

8月21日周五
  1. Hugging Face Blog65

    Liquid AI 发布 LFM2.5-DSpark 草稿模型,推理吞吐最高提升 3.2x

    Liquid AI 为 LFM2.5-1.2B-Instruct、LFM2.5-2.6B 和 LFM2.5-8B-A1B 三个模型发布 DSpark 草稿模型 checkpoint,通过投机解码在不改变输出质量的前提下加速推理,GPU 吞吐最高提升 3.18x,端侧最高 2.87x。

    推荐理由:官方为 LFM2.5 三款模型发布 DSpark 草稿模型,给出从 H100 到 MacBook 的实测加速数据和开源接入方式。

8月19日周三
8月18日周二
8月17日周一