跳到正文

推理能力

模型推理能力的进展:思维链、推理模型、数学与逻辑基准的突破与争议。

当前仅显示精选新闻

最新精选

第 41–60 条 · 共 207 条
9月17日周四
9月12日周六
  1. AI寒武纪 · 微信公众号76

    陶哲轩等25位菲尔兹奖得主发布联合声明,批评AI刷榜式解题破坏数学研究

    陶哲轩联合25位菲尔兹奖得主发布题为《数学领域中人工智能的严重失衡》的联合声明,批评AI公司把攻克数学难题当作基准测试来推进,认为这与数学界的目标严重脱节。声明指出,近几个月大语言模型的数学能力大幅提升,但AI主导的解题成果往往仓促发布,来不及严谨论证、提炼新方法和规范引用前人工作,引发成果归属与抄袭争议。声明认为这属于更广泛的人工智能对齐问题的一部分,2026年菲尔兹奖获得者邓煜也在签署人之列。

    推荐理由:声明全文与25位签署人名单完整呈现,读者可了解数学界对AI以解题跑分推进研究的具体担忧。

9月11日周五
9月10日周四
  1. @AravSrinivas66

    DeepSeek 发布新架构家族中最小模型 DeepSeek-V4.1-Flash,具备原生视觉理解,主打更强能力、更快推理、更高吞吐,并称可扩展至更大模型。官方以 1/6 系列推文介绍该模型,Perplexity CEO Aravind Srinivas 转发并配文 Wow。配图给出它在 Terminal-Bench 3.0、DeepSWE v1.1、CyberGym 和 Automation-Bench 上的对比数据。

    引用@deepseek_ai@deepseek_ai

    🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6

    推荐理由:DeepSeek 新架构家族最小模型亮相,配图给出它在四个基准上与其他模型的对比数据,可据此了解该型号的定位与表现。

9月9日周三
  1. IT Home78

    AI 攻克数学难题,纽约大学教授质疑 OpenAI「截胡」其研究成果

    纽约大学数学教授 Tristan Buckmaster 宣布三项证明成果,并质疑 OpenAI 在其成果公开前就基于其工作推进,抢先公布了纳维-斯托克斯存在性与光滑性问题的完整证明。

    推荐理由:原文给出了双方时间线与算力成本,读者可据此观察 AI 在数学研究中的成果归属与数据使用争议。

  2. @Thom_Wolf73

    Hugging Face 联创 Thomas Wolf 在 OpenAI 宣布用一组智能体给出 Navier-Stokes 问题解之后表示,这不足以说明 AI 已经解决数学。他指出这次是通过证明猜想为假(反例)而非完整证明取胜,近期 AI 数学的前沿成果也多是反例或大海捞针。他认为模型仍缺少数学家依赖的数学品味,难以区分优美论证与仅仅正确的论证,并希望这类工具的访问能被尽可能广泛共享。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:Hugging Face 联创从数学品味角度提出,AI 在数学上的前沿成果多是反例式搜索,尚不足以判定领域已被解决。

  3. @AYi_AInotes82

    OpenAI 发布公告称,一组 AI Agent 在 88 小时内给出了纳维-斯托克斯方程的证明,该问题已悬置约 90 年。公告称证明由一个能力显著强于 GPT-6 Astra 的下一代模型产出,GPT-6 Astra 用 17 小时完成 Lean 形式化代码核验。转发该公告的作者补充称,此次动用了约 10,000 个智能体、270 万条上下文与 1300 亿输出 Token。

    原始视频预览图;未保存可播放视频URL
    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 称一组智能体在 88 小时内给出纳维-斯托克斯方程证明,读者可据此了解大规模智能体协同与 Lean 形式化核验的用量。

  4. @SemiAnalysis_65

    OpenAI 称正在分享 Navier-Stokes 千禧年大奖难题的解法,该证明由一组智能体使用比 GPT-6 Astra 更强的下一代模型产出。问题涉及三维流体运动的光滑解是否会崩溃,已悬置约 90 年。SemiAnalysis 的推文本身只有两个链接。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 称一组智能体用比 GPT-6 Astra 更强的下一代模型给出 Navier-Stokes 问题的解法,可据此看智能体在前沿数学中的角色。

  5. @rohanpaul_ai77

    OpenAI 分享称,其内部智能体系统使用比 GPT-6 Astra 更强的下一代模型,提出对 Navier-Stokes 方程光滑三维流体运动是否会突然失效这一约 90 年未解问题的解,并同时公开证明文稿与 Lean 形式化版本。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:原文给出智能体搜索的耗时与 token 消耗,可了解大规模智能体协作做数学研究的具体形态。

  6. Noam Brown65

    Noam Brown 回应争议称,OpenAI 的 Navier-Stokes 成果并非依赖 Levent 或 Tristan 的提示词,没人查看过这些提示词。他附上图表显示,在一组公开数学题上 OpenAI 内部模型的 pass rate 在相同 test-time compute 下显著高于 GPT-6 Astra,以此说明所用模型较当今 LLM 有大幅提升。

    引用Sebastien Bubeck@SebastienBubeck

    I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.

    推荐理由:作者以本人身份回应 Navier-Stokes 归属争议,并用内部模型与 GPT-6 Astra 的对比图说明成果不依赖外部提示词。

  7. Eric70

    OpenAI 宣布给出纳维-斯托克斯千禧年大奖难题的一个解,证明由一组智能体使用比 GPT-6 Astra 能力更强的 OpenAI 下一代模型产出。该问题关注纳维-斯托克斯方程描述的光滑三维流体运动是否会崩溃,约 90 年来未获解决。转发作者以在 OpenAI 工作的口吻感叹又是一个疯狂的日子。

    引用OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 官方宣布用下一代模型的智能体群产出纳维-斯托克斯千年问题证明,读者可以关注智能体做数学研究的这一路径。

  8. @EMostaque69

    OpenAI 称一组智能体借助一个能力显著强于 GPT-6 Astra 的下一代模型,给出了 Navier-Stokes 千禧年难题的一个解;该问题关注光滑三维流体运动在方程下是否会失效,已悬置约 90 年。Emad Mostaque 转发并称这是通往 ASI 的标志性事件,同时表示找到一个解不代表没有其他解,其贴出的截图显示智能体在 NavierStokes 文件夹中检索文件并变更了 1 个文件。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:OpenAI 称一组智能体借助下一代模型给出 Navier-Stokes 方程解,可了解这一数学难题的 AI 攻关方式与外界反应。

  9. @omarsar072

    OpenAI 称一组智能体使用一个比 GPT-6 Astra 能力显著更强的下一代模型,给出了 Navier-Stokes 千禧年难题的证明。该问题关乎 Navier-Stokes 方程所描述的平滑三维流体运动是否会失效,已悬置约 90 年。转述此事的 @omarsar0 表示,若 GPT-6 Astra 已有这样的水平,下一代模型的表现可以想象。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:转述 OpenAI 关于千禧年数学难题证明的说法,读者可借此了解其下一代模型与智能体协作的定位。

  10. @kimmonismus69

    OpenAI 称其 AI 解出了约 90 年未解的 Navier-Stokes 千禧年难题,解由一个约 1 万个并发智能体的小组产出,所用内部模型能力明显强于 GPT-6 Astra。智能体 88 小时得到解,Lean 形式化与验证再由 GPT-6 Astra 花 17 小时完成,单是这一项就消耗 1300 亿输出 token。OpenAI 公开了证明与 Lean 形式化,内部模型仍在训练中。

    引用@OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

    推荐理由:材料给出了智能体攻关与 Lean 形式化验证的完整流程,读者可据此了解这一宣称结果如何被产出并接受机器核验。