“We find that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.” Eighteen months ago they scored well below human accountants Good discussion here: https://www.mercor.com/blog/human-baselines-for-benchmarks-ai-now-outperforms-junior-accountants/
#推理
#推理
今日 5 条
Rohan Paul@rohanpaul_aiAI 评分5959
引用Ethan Mollick@emollick
Thomas Wolf@Thom_WolfAI 评分4545


引用Bartosz Naskręcki@nasqretI cannot agree more. Kevin Buzzard made so many points I agree with. But the best one is this "I thus believe that in the future we will reach a new “natural boundary” in mathematics, beyond (and perhaps way beyond) where we are now, but where machines are going to get stuck and where it is not viable to expend any more resources to make the next big leap. (...) I believe that the optimal thing to do (...) is to let the machines loose, see what happens, and then begin the journey to where they have stopped." https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/
MIT Technology Review · AIAI 评分5555 AlphaGo 核心成员 Thore Graepel 撰文:LLM 并不会真正推理
前 DeepMind AlphaGo 团队核心成员、UCL 教授 Thore Graepel 撰文称,Move 37 靠的是搜索机制构成的推理而非纯直觉,而 LLM 的 next-token 预测与链式思考仍属系统 1。
Yuchen Jin@Yuchenj_UW精选AI 评分6767引用Andrej Karpathy@karpathyWe'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Latent SpaceAI 评分5252 Latent Space 访谈 MIT 的 Alex Zhang:RLM、harness 设计与研究品味
Latent Space 播客访谈 MIT 博士生 Alex Zhang,围绕其 Recursive Language Models(RLM)研究展开。
Anthropic Research精选AI 评分8282 Anthropic 评测 GLM-5.3:可自主构建端到端漏洞利用且防护易被绕过
Anthropic 发布对智谱 GLM-5.3 的网络安全能力分析,认为它是首个在无实质防护下开放权重的强网络攻击能力模型,与 NIST CAISI 评估结论大致一致。
推荐理由:Anthropic 以一手评测数据说明 GLM-5.3 的漏洞利用能力与防护绕过率,并解释攻击者可及性与 Claude 的访问限制差异。
François Chollet@fcholletAI 评分5353Tomer TunguzAI 评分5757 Tomer Tunguz 解析 GPU 租金翻倍至 $8.08 而推理价格仍在下降的原因
B200 GPU 租金九个月内从 $4.40 翻倍到 $8.08 每 GPU 小时,但 AI 价格仍在下降。作者归因于数据中心建设成本上升、需求爆发与推理效率提升并存:同一基准的完成成本从 $0.55 降到 $0.0015,Microsoft 称每 GPU 生成 token 数同比增长 90%。
elsewhere articles精选AI 评分6868 从 MiMo-V2.6 看大模型「斩杀线」:斩的是中间层模型
文章以小米 9 月开源的 MiMo-V2.6 系列为切入点,分析大模型「斩杀线」概念,即在智能和成本两个维度都被超过的模型会失去被选择的理由。
推荐理由:文章借小米 MiMo-V2.6 梳理了智能与成本双维度的行业竞争框架,读者可以据此理解 Agent 时代效率为何成为模型竞争的新坐标。
-Zho-@ZHO_ZHO_ZHOAI 评分3636
Nathan Lambert: InterconnectsAI 评分5353 Nathan Lambert 对谈 Epoch AI 的 JS Denain:RSI、中美差距与蒸馏
Interconnects 播客中,Nathan Lambert 与 Epoch AI Insights 负责人 Jean-Stanislas Denain 讨论多项议题。
Jeff Dean@JeffDeanAI 评分3939引用Dawn Song@dawnsongtweetsI had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8
elsewhere articles精选AI 评分6565 Sayash Kapoor 与 Arvind Narayanan 分析 AI 会不会让论文更快却让科学进步更慢
Sayash Kapoor 和 Arvind Narayanan 在 2025 年发表的文章提出生产—进步悖论:全球论文数量约每 12 年翻一番,1900 至 2015 年间增长约 500 倍,但诺奖成果诞生于获奖前 20 年内的比例从 1970 年约 90% 降至 2015 年约 50%,科学进步相对投入明显放缓。
推荐理由:文章把 AI 加速科研的讨论从模型能力转向注意力、激励、可复现性和人类理解等制度瓶颈,提供了评估 AI 科研工具的三个问题。
Noam Brown@polynoamialAI 评分5151引用Dwarkesh Patel@dwarkesh_spNew episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
Dwarkesh Patel精选AI 评分6969 Dwarkesh 对谈 OpenAI Noam Brown:Agent 集群、对齐与递归自我改进
Dwarkesh Patel 与 OpenAI 研究员 Noam Brown 对谈,涉及用 1 万个 AI Agent、1300 亿 token、88 小时求解 Navier-Stokes 千禧年大奖问题的工作。
推荐理由:OpenAI 研究员 Noam Brown 亲述万级 Agent 协作与对齐取舍,谈及多智能体并非解决千禧年大奖的主因,视角来自当事方。
Eric@ericmitchellaiAI 评分1414
Dwarkesh PatelAI 评分5656 Dwarkesh 对谈 John Schulman、Beren Millidge 与 Charlie O'Neill:递归自我改进还有多远
Dwarkesh Patel 与 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman、Baseten 模型训练负责人 Charlie O'Neill 长篇对谈,逐段讨论递归自我改进(RSI)最可能失败的技术原因、中国实验室的追赶路径、自动化 AI 研究者的训练方式以及长时程 RL 能否带来 AGI。
Eric@ericmitchellai精选AI 评分7070引用OpenAI@OpenAIWe’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
推荐理由:OpenAI 官方宣布用下一代模型的智能体群产出纳维-斯托克斯千年问题证明,读者可以关注智能体做数学研究的这一路径。
@emollick@emollickAI 评分3030
Dwarkesh PatelAI 评分5858 Dwarkesh Patel 提出持续学习时代的 8 项 AI 预测
Dwarkesh Patel 提出持续学习到来后的 8 项预测,认为模型仅在会话间写 Markdown 无法积累执行整份工作所需的经验,经验必须沉淀进权重。他据此推论部署前安全检查将失效、对齐技术需重构、领先实验室将靠部署数据加速拉开差距,并以 Anthropic 内部自 2 月起使用 Mythos、6 月才公开发布的 4 个月差距为例说明先发部署的学习优势。
elsewhere articlesAI 评分6161 十字路口对话生数科技张金涛:从 SageAttention 到 Vidu S1 的推理加速与实时交互视频
播客邀请生数科技张金涛聊实时交互视频模型 Vidu S1:上传一张照片可变成能实时视频通话的 AI 角色,生成速度超过播放速度并支持无限时长。
AI as Normal Technology精选AI 评分7272 Arvind Narayanan 在 ICML 演讲谈 AI 时代还剩什么工作
Princeton 的 Arvind Narayanan 在 ICML Seoul 发表题为“还有什么工作留给我们做”的主题演讲,主张用 AI as Normal Technology 框架看待AI影响,并称实验室里程碑不会突然让人失业。他提出方法、产品、早期采用、适应四阶段,指出可靠性指标两年内仅提升五到十个百分点、适应阶段需要数十年;未来工作将从构建转向评估,人类应与AI形成“共同超级智能”。
推荐理由:作者结合能力与可靠性的测量数据,把AI经济影响拆为四个阶段,给出职业适应与评估优先的判断框架。
Dwarkesh PatelAI 评分3232 Adam Brown 深入浅出讲解广义相对论
Google DeepMind BlueShift 负责人 Adam Brown 在播客中深入浅出讲解广义相对论,从爱因斯坦"最快乐的思想"切入,剖析引力是时空弯曲而非力,并延伸至黑洞为何无法被用来无限提取能量。节目最后讨论了 AI 距离从零重新发现广义相对论还有多远。
Import AIAI 评分5959 Import AI 464:Fable 写出 GPU kernel,AI 自动化与模拟计算
Jack Clark 在 Import AI 464 期中汇总多项进展:Fable 在 KernelBench-Mega 上提交了首个也是最快的 megakernel。
Dwarkesh PatelAI 评分4545 Grant Sanderson 谈 AI 与数学的未来
3Blue1Brown 的 Grant Sanderson 在播客中讨论 AI 在数学领域的进展。他指出 AI 自 2024 年起能在 19 秒内解出 IMO 几何题,但在组合数学题上仍表现吃力,数学内部的进展呈现"分形"式的不均衡。他还探讨了 AI 能否找到领域间隐藏联系、概念性突破的验证周期可能长达一个世纪,以及 AI 是否净增人类对数学的理解。
Dwarkesh Patel精选AI 评分6565 Dwarkesh Patel:下一个突破是 AI 在工作中学习
Dwarkesh Patel 撰文认为,实验室押注的 RLVR 训练未必能泛化到无法在数据中心内复现的现实领域,因为训练样本效率低且缺少可重放模拟器。
推荐理由:作者论证 RLVR 难以泛化到不可复现的现实领域,并梳理 OPSD 与模拟训练等让模型在工作中持续学习的可能路径。
Import AIAI 评分6767 Jack Clark 撰文判断 2028 年底前无人类参与 AI 研发概率超 60%
Jack Clark 在 Import AI 455 中判断,到 2028 年底出现无人类参与的 AI 研发(模型自主训练出后继版本)的概率超过 60%,2027 年约为 30%。
AI as Normal TechnologyAI 评分6262 Arvind Narayanan 考证 Moravec 悖论:一个从未被实证检验的经验法则
Arvind Narayanan 在其新文章和视频中逐条考证 Moravec 悖论(对人类难的任务对 AI 容易、反之亦然),指出它从未被实证检验,本质是 AI 社区认为值得研究的问题造成的筛选效应,其演化解释也很可疑。
AI as Normal Technology精选AI 评分6969 AI 进步是否在放缓:Arvind Narayanan 等人解析模型扩展与推理扩展之争
Arvind Narayanan、Benedikt Ströbl 和 Sayash Kapoor 撰文分析 OpenAI、Anthropic 和 Google Gemini 下一代模型遇阻后的叙事翻转,认为宣布模型扩展已死为时尚早,行业领袖的预测并不可靠。
推荐理由:作者用可核查的证据指出业内叙事反复翻转的成因,并区分能力提升与实际社会影响的薄弱关联。