跳到正文

推理能力

模型推理能力的进展:思维链、推理模型、数学与逻辑基准的突破与争议。

当前仅显示精选新闻

最新精选

第 1–20 条 · 共 207 条
今天10月6日周二
  1. Rohan Paul65

    Reflection AI 发布开源智能体模型 Beam,总参数 501B、激活 23B,宣称在编码和智能体任务上推进西方开源前沿,完整权重将于本月以 Apache 2.0 发布并附带 FP8 和 NVFP4 版本。

    引用Reflection@reflection_ai

    Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active. - Frontier reasoning efficiency - Advances the Western open frontier on coding & agentic tasks - Trained end-to-end from scratch Full weights release this month. Learn more about Beam: http://reflection.ai/beam

    推荐理由:作者在转发基础上补充了效率估算口径和 headline 图表外的对比数据,帮助读者更冷静地看待 Beam 对 GLM-5.2 的领先说法。

10月5日周一
  1. AI前线 · 微信公众号70

    OpenAI DevDay 推多 Agent 产品 Dots,Noam Brown 称万 Agent 解千禧年难题功劳多 Agent 不足 10%

    AI前线编译 Noam Brown 与 Dwarkesh Patel 的对谈。OpenAI 在 DevDay 2026 推出全天候智能体 Dots 和开放 Harness、多智能体控制能力的 Agents API。

    推荐理由:Noam Brown 对多智能体实际贡献、规模扩展效率与对齐难点的内部视角,为判断多 Agent 路线提供了难得的校准参考。

10月3日周六
  1. Alexandr Wang65

    Meta 宣布数学家与 Muse Spark 1.1 和 Muse Spark 1.2(Thinking Mode)通过常规 meta.ai 界面协作,解决六个公开数学难题并发布六篇论文。协作无自定义研究脚手架,每篇论文标注哪些段落主要由人类或 AI 起草,并由第二组数学家审阅,亦承认其他团队独立公布的同题解法。

    引用AI at Meta@AIatMeta

    Following gold-medal-level performance from our AI models across five competitions in mathematics, physics, and chemistry, we asked a harder question: can AI contribute when a problem is genuinely open and without an existing solution path? Over the past several months, mathematicians worked with Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the regular http://meta.ai chat interface, with no custom research scaffold, to find solutions to such problems. Our goal wasn't to mass-produce papers, but empower researchers. Every collaboration followed the same principles: mathematicians guided the research, a second group of mathematicians then reviewed the work, each paper marks which passages were drafted primarily by humans or AI, and each credits the prior research it builds on. Where other teams independently announced solutions to the same problems, we acknowledge their work as well. Today, we're sharing six papers from that collaboration. 🧵👇

    推荐理由:原文展示了模型在无现成解法的公开数学难题上与数学家协作产出六篇论文的案例,并提供人机分工标注原则。

10月2日周五
  1. Rohan Paul65

    NVIDIA 在论文 Staying on Task 中测试 7 个开源模型处理加法、排序等重复任务,发现 128K token 任务的平均准确率比 4K token 任务低 62.8%,最好的模型在最长任务中也只有 17.1% 做到每项全对。模型会看似理解任务却中途丢失位置,尤其在条目没有 ID 时;作者建议给每个条目编号、小批次处理并逐行检查输出。

    推荐理由:论文给出长任务可靠性下降的具体数字,并附带可迁移的做法,适合构建长流程智能体时参考。

  2. Rohan Paul68

    Google 发布论文 Cogentic,让多个 Gemini 智能体同时探索不同证明方向,由专用组件进行对抗式验证,并将已证结果保存供后续轮次使用。多数问题仅需约 100 次模型调用,5 个开放问题的结果均经人类专家独立确认,涉及在线学习、拍卖理论和机制设计。论文地址 arxiv.org/abs/2609.40324。

    推荐理由:原文给出了 Cogentic 多智能体证明发现系统的结构设计与验证结果,其中的严格审查与已证工作记录方法可迁移到长任务 Agent 设计。

  3. AI前线 · 微信公众号68

    前微软合伙人 Ramez Naam 万字长文质疑 RSI:回路强度仅为智能爆炸门槛的 10% 到 20%

    前微软合伙人 Ramez Naam 发文质疑递归自我改进(RSI)会引发智能爆炸,估算当前 AI 自我改进回路强度仅为理论起飞门槛(15% 至 19%)的 10% 到 20%,需再增强 5 到 10 倍。

    推荐理由:长文用机构内部数据和经济学模型逐层核算,说明当前 AI 自我改进回路离智能爆炸门槛的量化差距。

  4. Yuchen Jin67

    Yuchen Jin 转引 Andrej Karpathy 关于理解语言模型输出的建议,并表示希望 AI 能直接生成一段 Karpathy 风格的视频,但如今没有 AI 能做到。Karpathy 在引用内容中提出几种输出形式,包括让 LLM 用航空维护文档的受控语言规范 ASD-STE100 解释概念、生成图表和交互式 HTML 网页,以及用 ElevenLabs API key 或本地免费方案生成 3b1b 风格的讲解视频;他认为 LLM 会承担更多工作,人类的工作将上升为监督和理解。Yuchen Jin 还提到 Karpathy 已超过一年没有在 YouTube 上传视频。

    引用Andrej Karpathy@karpathy

    We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.

10月1日周四
  1. AGI Hunt · 微信公众号80

    谷歌 DeepMind 发布 Gemini 4 Argon,单次输出达 100 万 token 但编程偏科

    谷歌 DeepMind 发布新旗舰模型 Gemini 4 Argon,未带 Pro、Flash 后缀,单次输出上限从上一代 64K 提升到 100 万 token,介绍价每百万 token 输入 $2、输出 $10。

    推荐理由:作者汇总了官方跑分、第三方榜单和彭博社爆料,读者可以对照看到 Argon 各项能力的强项与短板。

  2. AI寒武纪 · 微信公众号77

    谷歌发布Gemini 4 Argon,单次输出上限提升至100万Token

    谷歌发布Gemini 4 Argon,单次输出Token上限从64K提升到100万。模型在DeepSWE v1.1取得77.9%的SOTA成绩,AutomationBench拿下51.3%头名,但Terminal-bench4.0仍落后。Argon百万输入Token收费2美元、输出10美元,目前通过Fairwind计划向受信任的网安人员开放,并参与美国政府发布前审查。

    推荐理由:文章汇总了模型能力、价格与内部落地数据,读者可以据此评估它在软件工程和网安场景的适用性。

  3. The Decoder79

    Google 发布 Gemini 4 Argon,追赶 OpenAI 与 Anthropic 但未取得明确领先

    Google 发布新旗舰模型 Gemini 4 Argon,是其七个多月来首款前沿模型,Artificial Analysis 测试中得 53 分,与 GPT-6 Astra (max)、Claude Fable 5.1 持平,但仍落后 Claude Opus 5.5 的 58 分。

    推荐理由:原文汇总了第三方测试与定价细节,指出 Gemini 4 Argon 缩小差距但未领先,且单价优势来自低 token 价格而非效率。

  4. InfoQ · 微信公众号71

    Google 发布 Gemini 4 Argon:输出上限提升至百万 Token,已参与 80 万行内核代码迁移

    Google 发布新一代模型 Gemini 4 Argon,面向长流程软件工程、企业知识工作与网络安全防御,输出 Token 上限从 64K 提升至 100 万,初期每百万输入 Token 2 美元、输出 10 美元。

    推荐理由:梳理了 Gemini 4 Argon 的跑分、定价与分阶段开放节奏,读者可以据此评估其能力与实际可用性之间的距离。

9月30日周三
9月29日周二