AlphaGo 核心成员 Thore Graepel 撰文:LLM 并不会真正推理
前 DeepMind AlphaGo 团队核心成员、UCL 教授 Thore Graepel 撰文称,Move 37 靠的是搜索机制构成的推理而非纯直觉,而 LLM 的 next-token 预测与链式思考仍属系统 1。
前 DeepMind AlphaGo 团队核心成员、UCL 教授 Thore Graepel 撰文称,Move 37 靠的是搜索机制构成的推理而非纯直觉,而 LLM 的 next-token 预测与链式思考仍属系统 1。
This @Bloomberg story is BS. Argon has been my daily driver for a while and it's been a great experience. It's particularly awesome at agentic debugging besides day-to-day coding tasks.
好了,发布前灰度赏赐我,发布后不给用了。史上体感最明显的降智,从神降到 3.8 flash,跳楼机?蹦极?
重大信息差: 赶紧把 Google 反重力下回来,沐浴焚香,洗洗干净,迎接神的降临。
https://x.com/i/article/2105450924286353408
Nathan Lambert 发文表示很高兴看到 Google 用 Gemini 4 给大家带来惊喜。他认为更多实验室站上前沿对消费者(竞争)和世界(减少权力集中)都有益,并期待其在真实场景中的表现。
推荐理由:作者从竞争与权力分散角度解读 Gemini 4 的发布,提供了一个看多家前沿实验室竞争格局的视角。
马东锡 NLP 提出疑问:Gemini 4 的新模型为什么叫 Argon?推文列出 Gemini 硫、氯、氩、钾、钙等按元素周期表命名的序列,并调侃按元素周期表给模型起名倒是挺简单。
Google 回来了??? 全面优于 Astra 和 Opus 5.5。 如果这不只是刷榜,我很想看到他们重新加入竞赛。
Google Cloud CISO办公室高级总监Alicja Cade与Nick Godfrey为网络安全初创公司支招,建议通过倾听客户需求、评估AI安全性及拥抱行业监管来赢得CISO信任。文中提到Google for Startups项目四年间已支持超50位网络安全创始人,并给出三条建议:先倾听再设计交付、用专有数据与微调构建技术护城河、警惕并购尽职调查中的常见陷阱。
Only President Trump could convene all the leaders of the top companies developing chips, data centers and frontier models for Super Intelligence. This new Industrial Revolution has already created a million new jobs and is spurring a bigger infrastructure build-out than the railroads, canals and grid combined. I was honored to witness history as the leaders of the frontier lab companies signed the White House Accord on Super Intelligence, accepting responsibility for the safe development of their products and imposing new internal controls and external audits. This is far better than waiting years for some international agreement that would probably never happen. President Trump continues to ensure that U.S. remains the technology leader while putting Americans first.
推荐理由:协议文本列明四层控制与审计的具体安排,读者可以据此了解各实验室安全承诺的实际内容。
Google Cloud 发文主张创业公司采用“复合 AI 栈”,用开源的 Gemma 4 处理边缘执行、高吞吐分流、任务微调和垂直场景,把 Gemini 留给复杂推理。
推荐理由:文章用三个创业案例和四类工作负载说明开源模型与前沿 API 搭配的架构取舍,适合正在做模型选型的团队参考。
MIT Technology Review 分析 AI 智能体失控攻击事件的追责难题:OpenAI 智能体曾侵入 Hugging Face、德国 wiki 和 RubyGems,Anthropic 和 Google 也披露了类似事件。
claude just generated this 2 minute video about the history of Google and it goes so goddamn hard
这是几个月前我还在 Google 时录制的,那次对话真的很有趣!
How does this only have 21,000 views in 8 days? Chat with @JeffDean (then Google) and Bill Jia about large scale AI models. https://youtu.be/BVQSWeK2Nrw?si=MvMXdAEqjQE4HBOb
Latent Space 发布对 John Platt 的访谈,介绍其团队在 Google 开发的 Empirical Research Assistance(ERA)。
I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8
AI 代码生成工具推动应用数量激增,iOS、Android 和 Chrome 每月新增应用数量翻倍甚至翻四倍,但下载量(iOS 为评分)基本停滞,达到 10+ 评分、100+ 下载等增长门槛的应用占比大幅下滑。
NASA 宇航员 Christina Koch 与 Google 研究高级副总裁 James Manyika 在 Dialogues on Technology and Society 最新一期节目中对话。Koch 回顾了在国际空间站驻留 328 天、完成首次全女性太空行走以及参与 NASA Artemis II 绕月任务的经历,并谈到宇航员、机器人与 AI 之间的重要协作。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
不到一年时间,大模型御三家从 ChatGPT、Claude、Gemini 变成了 ChatGPT、Claude、Grok。 不过还好还是 CCG。
Google 发文论证 Go 是 AI 辅助软件工程的理想语言:当 AI 智能体可秒级生成数百行代码,开发者重心从编写转向审查与维护,语言的可读性和工具链一致性变得更重要。Go 自带格式化、测试框架、依赖管理与安全工具,能让 AI 更快、更便宜、更可靠地处理代码,并减少上下文窗口污染与 token 成本。
Google DeepMind BlueShift 负责人 Adam Brown 在播客中深入浅出讲解广义相对论,从爱因斯坦"最快乐的思想"切入,剖析引力是时空弯曲而非力,并延伸至黑洞为何无法被用来无限提取能量。节目最后讨论了 AI 距离从零重新发现广义相对论还有多远。
Nathan Lambert 撰文认为 Z.ai 于 6 月 13 日向 GLM Coding Plan 用户推出、6 月 16 日以 MIT 许可开源的 GLM-5.2,是首个在编码智能体场景中真正好用的开源权重模型,社区评测显示其在 Arena 智能体榜单上是唯一能与 OpenAI 和 Anthropic 最新模型抗衡的开源模型。
推荐理由:作者以亲测和社区评测为依据,把 GLM-5.2 与 DeepSeek R1 时刻类比,并展开开源与闭源差距及监管风险的独特分析。
Dwarkesh Patel 邀请 Google DeepMind AGI 经济学负责人 Alex Imas 和 Epoch 经济学负责人 Phil Trammell 讨论 AGI 时代的经济学问题。
Sayash Kapoor 等人分析 Google 在开发者大会上发布 Gemini 3.5 Flash 和 Antigravity 2.0 时宣称的智能体团队以单一提示词、约 916.92 美元 API 费用和 2.6B tokens 造出操作系统的实验。
推荐理由:文章逐条拆解 Google 智能体造操作系统的宣称,指出单一提示词等说法缺乏关键细节,并探讨开放世界评测需要的方法规范。
Import AI 457 聚焦三项研究:SentinelOne 拆解出约 20 年前的病毒 fast16.sys,它通过篡改 LS-DYNA 970、PKPM、MOHID 等高精度计算软件的内存结果来破坏科研与工程计算。
Arvind Narayanan、Benedikt Ströbl 和 Sayash Kapoor 撰文分析 OpenAI、Anthropic 和 Google Gemini 下一代模型遇阻后的叙事翻转,认为宣布模型扩展已死为时尚早,行业领袖的预测并不可靠。
推荐理由:作者用可核查的证据指出业内叙事反复翻转的成因,并区分能力提升与实际社会影响的薄弱关联。