跳到正文

#推理

今日 17 条
今天10月2日周五
  1. Rohan Paul59

    Rohan Paul 评论称规则型工作不再是职业而是一条提示词,人类在规则型职业中成了慢、贵、易错的一方。其引用的研究显示前沿模型在中长度、定义清晰的会计任务上已快于且更准于初级会计师,Claude Opus 5 在任务上 20/20 全对且数分钟完成,而 12 名持证 CPA 得分在 0% 到约 90% 之间,多人未能在 3 小时内完成;图中还显示每达成一条评分标准 Claude 成本 $0.21,无 AI 会计师为 $10.35。

    引用Ethan Mollick@emollick

    “We find that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.” Eighteen months ago they scored well below human accountants Good discussion here: https://www.mercor.com/blog/human-baselines-for-benchmarks-ai-now-outperforms-junior-accountants/

  2. Hugging Face Daily Papers43

    视频大模型时序推理为何在输出层“褪色”:TAI 方法无需训练即可增强时序表征

    视频大语言模型(VideoLLMs)的时序推理能力在中间层达到峰值,却随层数加深逐渐衰减至输出层,导致反转帧序后预测结果往往不变。研究者据此提出 Temporal Activation Injection(TAI),在峰值层提取时序表征并注入后续层,无需训练即可在三个 VideoLLM 和四个基准上稳定提升时序推理,且对非时序任务影响极小。该研究已被 NeurIPS 2026 接收。

  3. Thomas Wolf45

    Kevin Buzzard(IMO 满分、数论学家、Lean 形式化数学先驱)写了一篇非常深刻的文章。 如果数学不只是关于“人类理解”,那它又关乎什么? 如果 AI 能力持续指数级增长,而“数学是无限的”,那会发生什么?

    引用Bartosz Naskręcki@nasqret

    I cannot agree more. Kevin Buzzard made so many points I agree with. But the best one is this "I thus believe that in the future we will reach a new “natural boundary” in mathematics, beyond (and perhaps way beyond) where we are now, but where machines are going to get stuck and where it is not viable to expend any more resources to make the next big leap. (...) I believe that the optimal thing to do (...) is to let the machines loose, see what happens, and then begin the journey to where they have stopped." https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/

  4. Chubby♨️45

    webAI 发布 3.66B 参数形式逻辑模型 TwIL-LM3-Pro,可在笔记本本地运行。其综合逻辑评测与 Qwen3-8B 持平,参数量不足后者一半,并在全部六项形式逻辑任务上领先 VibeThinker-3B。该模型基于 IBM Granite 4.2 后训练,Q4 GGUF 权重仅 2.09 GiB,可通过 llama.cpp 本地推理。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  5. TechCrunch · AI63

    Google 发布 Gemini 4 Argon,称其为迄今最强模型

    Google(Alphabet)发布新模型 Gemini 4 Argon,主打防御性网络安全,称其可自主发现、验证并修复关键软件漏洞,目前仅通过 Fairwind 安全计划向部分网络安全合作伙伴开放。该模型也用于编码、调试和代码库迁移等日常工程工作,并称在多项基准上显著领先 GPT-6 Astra 与 Anthropic 的 Fable 和 Opus。

  6. Hugging Face Daily Papers35

    HC-DLM:分层连续扩散语言模型

    研究者提出分层连续扩散语言模型(HC-DLM),将离散 token 生成与连续隐变量轨迹耦合在单一去噪过程中,训练目标由 token 似然的变分下界推导而来。该模型以隐变量作为唯一持久生成状态,每步从中读出 token 并反馈作为下一步隐变量更新的脚手架。在 Sudoku、Countdown 和 LM1B 上,同等模型规模下 HC-DLM 在谜题准确率和生成困惑度上均优于离散与连续扩散基线。

  7. Yuchen Jin67

    Yuchen Jin 转引 Andrej Karpathy 关于理解语言模型输出的建议,并表示希望 AI 能直接生成一段 Karpathy 风格的视频,但如今没有 AI 能做到。Karpathy 在引用内容中提出几种输出形式,包括让 LLM 用航空维护文档的受控语言规范 ASD-STE100 解释概念、生成图表和交互式 HTML 网页,以及用 ElevenLabs API key 或本地免费方案生成 3b1b 风格的讲解视频;他认为 LLM 会承担更多工作,人类的工作将上升为监督和理解。Yuchen Jin 还提到 Karpathy 已超过一年没有在 YouTube 上传视频。

    引用Andrej Karpathy@karpathy

    We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.

  8. Chubby♨️42

    来了:大量用户现在报告他们的查询正被路由到 Fable 5.5! - 不带网页搜索的查询现在能返回最新结果。 - 初步 SVG 测试远优于 Fable 5.1。 这只是时间问题。准备好,世界上最好的模型即将发布。

    引用Chetaslua@chetaslua

    🚨 Fable 5.5 is auto routing on web , this is the screenshot it edited for x without even prompted he knows tibo check https://claude.ai see if you are getting routed or not

  9. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  10. Ars Technica · AI68

    AI 系统以 16 块 GPU 击败史上最强 Stratego 选手

    来自 CMU、MIT、NYU 和 Stanford 的团队开发 AI 系统Ataraxos,以 15 胜 1 负 4 平击败公认史上最强 Stratego 选手 Pim Niemeijer,训练仅用 16 块 GPU 和几千美元。Stratego 是隐藏信息量大且时间跨度长的非完全信息游戏,此前连 DeepMind 也没能造出稳定战胜顶尖人类的机器。

10月1日周四
  1. The Decoder79

    Google 发布 Gemini 4 Argon,追赶 OpenAI 与 Anthropic 但未取得明确领先

    Google 发布新旗舰模型 Gemini 4 Argon,是其七个多月来首款前沿模型,Artificial Analysis 测试中得 53 分,与 GPT-6 Astra (max)、Claude Fable 5.1 持平,但仍落后 Claude Opus 5.5 的 58 分。

    推荐理由:原文汇总了第三方测试与定价细节,指出 Gemini 4 Argon 缩小差距但未领先,且单价优势来自低 token 价格而非效率。

  2. Karina52

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,面向编码、企业知识工作和网络安全防御等复杂工作流,即日起通过 Fairwind Program 向部分受信任测试者开放。作者引述其 PostTrainBench 得分 45.3%,高于 Gemini 3.1 Pro 的 21.99% 和 GPT-6 Astra 的 44.3%;评测表还显示其在自动化与智能体编码等多项基准领先,但在 FrontierSWE v2、Terminal-Bench 4.0 等项落后于对比模型。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

  3. fofr33

    非常激动地分享,Gemini 4 Argon 即将到来。迫不及待想尽快跟大家分享更多内容。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  4. The Verge · AI73

    Google 发布 Gemini 4 Argon,初期仅限受信任的网络防御者使用

    Google 发布新一代前沿模型 Gemini 4 Argon,称其在软件工程、法律金融等企业知识工作和网络安全防御方面具有前沿性能。初期仅向一组受信任的网络防御者开放,Google 正参与美国政府预发布模型访问的自愿流程并逐步扩大访问。模型已用于 Google 内部工作流,如大规模代码库迁移;Google 将在更广泛发布前加强防范滥用和提示词注入攻击、监测错位等安全措施。

9月30日周三
  1. Hugging Face Daily Papers43

    自回归 Transformer 如何从局部观测外推混沌系统的全局动力学

    小型自回归 Transformer 仅用受限参数区间的轨迹训练,就能在训练分布之外的闭环评估中复现倍周期分岔、混沌动力学和吸引子结构。在 logistic 映射上,模型复现了直至周期 128 的连续倍周期分岔,得到有限阶标度比 4.6687,与 Feigenbaum 常数误差在 5×10⁻⁴ 以内。研究还通过因果干预揭示了控制参数信息经注意力机制影响状态预测与闭环动力学的路径。

  2. Hugging Face Daily Papers35

    邻近监督更优:Neighborhood OPSD 自蒸馏方法提升数学推理模型性能

    Neighborhood OPSD(N-OPSD)通过局部参数扰动构建冻结专家池,将参考对齐修正转化为学生可用的监督信号,在 AIME 2024、AIME 2025 和 HMMT February 2025 三项基准上,将 Average@12 较标准 OPSD 分别提升 2.75、1.67 和 1.94 分(对应 Qwen3-1.7B、4B、8B)。推理时仅使用蒸馏后的学生模型。

  3. Hugging Face Daily Papers32

    SpatialCORE:让大型视觉语言模型基于置信度进行空间推理

    SpatialCORE 是一个后训练框架,将模型对生成式 grounding 的自身置信度转化为空间推理的学习信号,通过自调节空间奖励按坐标 token 置信度加权每个预测边界框的匹配质量,并用答案门控将 grounding 优化与最终答案正确性绑定。该框架在多个基准上取得开源及专用空间推理模型中的 SOTA 结果,并可零样本迁移到未见数据分布,源代码已公开。

  4. Hugging Face Daily Papers41

    通过位置选择性自蒸馏从语言反馈中训练 LLM 评判模型

    研究提出位置选择性自蒸馏方法,利用自然语言反馈训练 LLM 评判模型。该方法基于教师与学生模型间的逐位置熵偏移进行位置掩码,保留熵偏移分布的低尾部分,从而提升分布外泛化能力。实验显示,自蒸馏评判模型在主观任务子类别上比 GRPO 等结果监督强化学习训练的评判模型高出 2-9 个百分点,在客观任务上保持竞争力。

9月29日周二
  1. Hugging Face Daily Papers36

    FlexRouter:为灵活 LLM 路由学习互补模型集合

    FlexRouter 是一个显式建模模型互补性的 LLM 路由框架,以「答案覆盖」为目标,最大化所选模型中至少一个给出正确答案的概率。它用 Determinantal Point Processes(DPPs)建模路由策略,并通过基于失败集边缘化的训练目标直接优化覆盖,推理时采用边际对数行列式增益的贪心策略,无需预设预算即可自适应确定子集大小。