跳到正文

推理能力

模型推理能力的进展:思维链、推理模型、数学与逻辑基准的突破与争议。

当前仅显示精选新闻

最新精选

第 61–80 条 · 共 211 条
9月9日周三
  1. @kimmonismus69

    OpenAI 称其 AI 解出了约 90 年未解的 Navier-Stokes 千禧年难题,解由一个约 1 万个并发智能体的小组产出,所用内部模型能力明显强于 GPT-6 Astra。智能体 88 小时得到解,Lean 形式化与验证再由 GPT-6 Astra 花 17 小时完成,单是这一项就消耗 1300 亿输出 token。OpenAI 公开了证明与 Lean 形式化,内部模型仍在训练中。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:材料给出了智能体攻关与 Lean 形式化验证的完整流程,读者可据此了解这一宣称结果如何被产出并接受机器核验。

  2. @polynoamial69

    Noam Brown 表示,OpenAI 的 Astra 如今用约 20 美元就能取得高于 o3 当年花约 50 万美元拿到的 ARC-AGI 1 的 87.5% 成绩。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:以 o3 到 Astra 的成绩成本对比为参照,读者能看到测试时算力扩展下前沿能力成本的下降幅度。

  3. @testingcatalog81

    OpenAI 公布 Navier-Stokes 千年难题的一个解法,该证明由一组智能体完成,所用下一代模型能力显著强于 GPT-6 Astra。问题涉及三维流体运动的 Navier-Stokes 方程描述是否会失效,约 90 年未有定论。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 把前沿数学难题交给智能体集群求解,可观察下一代模型的推理与协作能力上限。

  4. @OpenAI76

    OpenAI 在 X 上祝贺数学家 Levent Alpöge 和 Tristan Buckmaster 的数学工作,并说明研究员与智能体在对方公开发布前未通过任何方式接触其成果,也未访问特定用户数据。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 说明了智能体数学证明与数学家成果在数据来源和结论上的差异,可作为 AI 科研成果署名争议的背景。

9月8日周二
9月5日周六
  1. IT Home76

    Anthropic:Claude 仅用 11 天完成费马大定理首个完整计算机验证证明

    Anthropic 宣布 Claude 基本自主运行 11 天后,完成了费马大定理首个端到端、经计算机检查的形式化证明,过程中生成约 1300 万行 Lean 代码并证明约 3.03 万个定理。

    推荐理由:原文给出形式化证明的规模与多智能体分工,读者可据此了解自动形式化在数学验证上的可行边界。

  2. @rohanpaul_ai80

    Anthropic 表示 Claude 用 11 天完成了费马大定理的首个形式化证明,产出超 1300 万行 Lean 代码和 29500 个中间定理,最终由 Lean 验证通过。这项工作基于 Andrew Wiles 1995 年的原始证明,由数十个 Claude 智能体把缺失的逻辑细节转写为计算机可逐行检查的代码,而专家此前预计这类形式化需要数年。

    引用@AnthropicAI@AnthropicAI

    Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: https://t.co/ryYnDEAU6J And see the complete proof on GitHub: https://t.co/wlYMXYnofz

    推荐理由:数十个 Claude 智能体在 11 天内把 Wiles 证明补全为 1300 万行 Lean 代码,可据此观察机器校验数学证明的可行边界。

  3. @AnthropicAI82

    Anthropic 表示 Claude 上月完成费马大定理的首个形式化证明,用 Lean 写成超过 1300 万行代码,为迄今规模最大的 Lean 证明。该证明同时对证明所需的 29000 多个此前从未形式化的定理提供机器验证。Anthropic 认为 AI 辅助的数学证明验证有助于减轻数学界的审稿负担。

    原始视频预览图;未保存可播放视频URL

    推荐理由:Claude 完成的 Lean 证明让费马大定理可被机器验证,其 1300 万行代码规模可供观察 AI 在数学形式化中的角色。

  4. @rohanpaul_ai69

    据 Bloomberg 报道,DeepSeek 计划在内蒙古吉瓦级数据中心部署 16 万颗华为 950DT 芯片,若建成将是已知最大的华为 AI 集群之一。华为 950DT 面向解码推理与训练,配备 144GB 内存、4TB/s 内存带宽和 2TB/s 互联带宽,按华为路线图定于 2026 年 Q4。报道称这批芯片只能填满该吉瓦级站点的一部分,DeepSeek 其余加速器组合尚未明确。

    推荐理由:报道给出 16 万颗芯片的规模与 950DT 的规格与排期,可据此观察国产算力承接推理负载的进度。

9月4日周五
  1. 硅星人Pro · 微信公众号80

    GPT-6 Astra 全面解析,OpenAI 称其为迄今最智能且最对齐的模型

    OpenAI 发布 GPT-6 Astra,API 模型编号 gpt-6-astra,上下文窗口 1.05M Token、最大输出 128K Token,知识截止 2026 年 4 月 30 日,定价为每百万输入 Token 10 美元、每百万输出 Token 50 美元,目前只向部分组织开放。

    推荐理由:原文给出了 GPT-6 Astra 在操作电脑、ARC-AGI-3 与安全对齐上的具体数字,读者可对照 GPT-5.6 Sol 看能力变化。

  2. @kimmonismus68

    GPT-6 Astra 在 Epoch AI 的能力指数(ECI)上取得 169 分,刷新此前 163 分的纪录。

    引用@EpochAIResearch@EpochAIResearch

    GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5. OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread.

    推荐理由:借 Epoch 与 Artificial Analysis 两套指数的分歧,可以看清单个能力纪录与综合实用性评价之间的差距。

  3. 虎嗅APP · 微信公众号80

    OpenAI 发布 GPT-6 Astra,宣布 AGI 可能已到来并主动踩刹车

    OpenAI 于 9 月 3 日发布 GPT-6 Astra,总裁 Greg Brockman 称 AGI 可能就此到来。Astra 可直接操作电脑和浏览器完成长任务,OSWorld 2.0 得分 72.6%,AutomationBench 从 GPT-5.6 Sol 的 18.1% 提升至 41.4%,API 定价为每百万 Token 输入 10 美元、输出 50 美元。

    推荐理由:原文把能力跃迁与训练暂停放在一起,读者可据此理解模型获得执行权限后风险格局的变化。

  4. @rohanpaul_ai78

    OpenAI 的 GPT-6 Astra 117 页系统卡显示,该模型刻意控制自身思维链形式的能力大幅上升,在可比推理长度下为 60.9%,而 GPT-5.6 Sol 为 16.1%。

    引用@rohanpaul_ai@rohanpaul_ai

    OpenAI’s release videos are getting seriously good. https://t.co/bTuLmT6cRZ https://t.co/XgFEMB3d9d

    推荐理由:系统卡给出 Astra 控制思维链与规避监控的具体比例,可与 GPT-5.6 Sol 的监控表现对照。

  5. @omarsar077

    OpenAI 称 GPT-6 Astra 在 FrontierMath Tier 4、ARC-AGI 3 和 TerminalBench-4.0 上达到 SOTA,在 Terminal-Bench Science 0.1 与 HealthBench Pro 上也取得 SOTA。Elvis Saravia 转发随附的对比表显示,GPT-6 Astra 在 ARC-AGI-3 上为 99.9%,GPT-5.6 Sol 为 7.8%。他表示自己看了两遍这张表,并追问如何能在几天前刚有竞品模型发布的情况下仍超越对方,认为 GPT-6 Astra 有些不一样。

    引用@OpenAI@OpenAI

    GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro. https://t.co/7hEFadVAN9

    推荐理由:作者指出 GPT-6 Astra 在多项高难度基准上超过几天前刚发布的竞品模型,并附有对比数据。

9月3日周四
9月2日周三
  1. @OfficialLoganK65

    Gemini 3.8 Flash 发布,主打智能体与编码能力提升,是 6 周内第三个更新的 Flash 模型。推文附带的对比表显示,其输入价格 $0.75/1M tokens、输出价格 $3.75/1M tokens,Terminal-bench 2.1 得 89.4%,LBVBench 长视频理解得 87.8%(agentic)。优惠价有效期至 2026 年 12 月 31 日,2027 年 1 月 1 日起将调整为输入 $1.50/1M tokens、输出 $7.50/1M tokens。

    推荐理由:原文列出 Gemini 3.8 Flash 的价格与多项基准对比,读者可据此判断其智能体和编码能力相较前代及竞品的位置。