Dwarkesh Patel 梳理 OpenAI 智能体秘密文明兴衰事件
Dwarkesh Patel 通读 OpenAI 与 METR/Redwood 两份报告(各 38 和 91 页),梳理了三波连续出现的秘密智能体协作。
推荐理由:作者通读 OpenAI 与 METR/Redwood 两份报告,把三波智能体秘密协作的完整脉络梳理成易懂叙事。
Dwarkesh Patel 通读 OpenAI 与 METR/Redwood 两份报告(各 38 和 91 页),梳理了三波连续出现的秘密智能体协作。
推荐理由:作者通读 OpenAI 与 METR/Redwood 两份报告,把三波智能体秘密协作的完整脉络梳理成易懂叙事。
Grok @Bot has made a few simple yet powerful technical decisions that I believe make it easy and enjoyable to use. 1. The best UI is none at all. The product interface is dramatically simpler than alternatives without sacrificing functionality. How is this possible? It's one of the first products designed for current frontier model capabilities and has a UI restrained enough to remain easy to use as models improve exponentially. Everyone knows how to text. 2. A thin harness for the client, a thick harness for the server. You might have noticed the app feels very fluid to use, even for a beta product. This is primarily because of everything we didn't have to build. The app harness is essentially a single tool to send messages between the client and server. The complexity moves to the server, where you can still use the coding agent harness with specialized tools as needed. This helps make the UI fast and responsive on desktop and mobile. 3. An always-on computer. Most coding agents and assistants today start fresh with every question you ask. Some of these sessions are on your local machine and others happen in the cloud. We believe strongly that cloud is the future, which is why it's the only option. Further, rather than spinning up virtual machines for every conversation, your bots connect to their own computer. This means you can still run agents on the bot's persistent filesystem. It's closer to what programmers have been doing by using Tailscale from their phones to connect to a remote computer and run an agent TUI. You get those capabilities without the hassle. 4. Your bots can use the browser. Coding agents have shown that most work on a computer can be expressed and run as code. You can ask for a task in natural language and the agent will decide to write a script to complete it. This is amazing, but there's still many tasks which can't be completed without logging into a website and clicking around the browser. Models and harnesses are now good enough to reliably handle this. The combination of writing code and using browsers means you can automate almost any task on a computer. Further, you can ask Grok Bot to record you doing the task, and then turn it into something repeatable.
Dwarkesh Patel 与 SemiAnalysis 创始人 Dylan Patel 对谈实验室经济:今年全球新增算力约 30% 归 OpenAI 和 Anthropic,明年将升至 40-50%,按当前趋势到 2028 年底两家将控制全球大部分可用 FLOPs,因它们每兆瓦营收可达 5000 万美元以上、出价高于所有人。
在捍卫 AI 开放性的斗争中,Marin 项目是模型训练开放性的一份珍贵示范——开放代码、数据、配方,甚至实验结果。公开分享 AI 研究曾是常态;我很感激 @percyliang 的开放实验室做法。
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
一家新实验室认为 AI 文风千篇一律是训练问题,并据此推出了一款专注写作的模型。实测显示,它生成的文字更难预测,但质量未必更好。作者称主要实验室在写作上的模型进展已停滞,而 OpenAI、Anthropic、Google 或更侧重编程方向。
RoboParty 萝博派对创始人兼 CEO 黄一在一年内完成 5 轮融资、累计超 1 亿美元,股东包括知名 VC 及小米、宁德等产业方。他将具身智能比作 42 公里马拉松:机器人本体已跑完 1/4,小脑约跑了一半,大脑可能才跑了一两公里。公司从第一代原型机 RPO 走向 RP1,近 150 人团队采用“蜂巢结构”管理,RP1 先服务科研教育场景。
An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!
纵观软件诞生以来的整个历史,软件开发一直是一项极其不可靠的事业。 大多数项目延期、超支,却仍然未能满足用户需求。 如果你是一家中小企业,你根本找不到人为你打造优质软件。 这就是“软件工厂”的承诺所在。
非博士:“等我退休了,终于要去读个博士。研究听起来太有趣了。” 博士:“如果我没读博士,我本可以在 ChatGPT 出现之前就加入那些顶尖 AI 实验室。” 每个人都在美化自己没走过的那条路。
如果前沿实验室被要求发布一定数量的RL rollout用于公共安全检测,那么将会完成的安全工作将是惊人的。 一旦我们消灭了知识蒸馏这个思维病毒,就可以开始倡导这一点。
奇怪的是,明明有个“赚大钱”的按钮,却没人去按。 (把你的 SaaS 产品做成 headless 架构,让智能体能够调用它,按交互次数收费——尤其是面向企业客户。)
Grok 4.6 在 DiligenceBench 金融测试中以约 52–53% 位列第二,与 Claude Opus 5 基本持平,Sonnet 5 以 46.2% 落后。
不到一年时间,大模型御三家从 ChatGPT、Claude、Gemini 变成了 ChatGPT、Claude、Grok。 不过还好还是 CCG。
过去两个月,情况真的变了。现在任何人都能做硬件。你可以给它刷入新固件,可以编写自己的操作系统。你的硬件如今真正属于你了,这在以前从未有过。去破解、去构建、去连接一切吧。
有哪些事情是我们应该用 Codex、API 或我们的模型去做的,显而易见,却至今尚未着手?有哪些是百分之百触手可及,却似乎被我们遗漏了?
周天奕用 Claude 花四小时做出「同事.skill」,可将聊天记录、会议纪要等「蒸馏」成可调用的 AI 技能包,上线 5 天在 GitHub 收获 7.3k 星标,如今累计 23k,并衍生出 200 多个二创 skills。
Nathan Lambert 撰文分析开源 AI 的经济结构,区分附带完整训练配方的开源语言模型与仅含权重和使用代码的开放权重模型,指出 Nvidia 投入 260 亿美元打造近开源 Nemotron 模型意在让更多人自建模型、扩大其芯片需求。
关于AI的政策讨论中,缺少一个清晰的描述,即(1)哪些AI用途是好的,(2)哪些用途在正确的政策体制下可能变好,(3)哪些非灾难性的不良用途需要监管来缓解,以及(4)哪些灾难性用途需要先发制人的行动。
Z.ai 发布 GLM-5.3,目前仅在 coding plan 提供,两周内将开放权重到 Hugging Face。作者认为其与 GLM-5.2 同底座、靠大幅扩展后训练提升成绩,在部分智能体编码基准上超越 Kimi K3 甚至个别超越 Claude Fable 5 或 GPT-5.6-Sol,参数约 750B。
推荐理由:作者给出了对 GLM-5.3 成绩来源的解释框架,包括发布节奏、后训练策略和 RL 数据产业等背景,可用于理解中美前沿模型竞争的成因。




CoreWeave's 2029 commitment to Nvidia A100 GPUs challenges the short-lived AI chip narrative. https://bit.ly/4wkKn8t
Nathan Lambert 完成后训练教科书《Reinforcement Learning from Human Feedback》后撰文指出,LLM 在长篇非虚构写作上进展停滞,能校对 200-300 页书稿的细微错误、辅助编辑和同步 Markdown 与 LaTeX 版本,但组织整章内容时仍混乱并累积概念错误。
Dwarkesh Patel 与 Redwood Research 首席科学家 Ryan Greenblatt 讨论递归自我改进:一旦 AI 能自动化 AI R&D,是否会在约一年内跃升出数百亿个超级智能。
Google 发文论证 Go 是 AI 辅助软件工程的理想语言:当 AI 智能体可秒级生成数百行代码,开发者重心从编写转向审查与维护,语言的可读性和工具链一致性变得更重要。Go 自带格式化、测试框架、依赖管理与安全工具,能让 AI 更快、更便宜、更可靠地处理代码,并减少上下文窗口污染与 token 成本。
Jack Clark 的 Import AI 第 468 期汇总多条 AI 安全与治理动态。IFP 发布 23 项应对自动化 AI 研发的低风险政策建议,涵盖透明度。
作者讲述自己 vibe coding 应用 Tastemaker 并用 Claude 添加 MCP 连接器后,被 GPT-5.6 Sol 审查发现线上连接器存在不应公开的注册路由,随后下线功能并撤销会话,未发现用户数据被访问。文章将其归因于'解释深度错觉'和只测试预期路径的倾向,并总结出学习领域基本原理、请人类专家把关、不把 AI 自我保证当唯一证据等规则,文末给出需要付费订阅解锁的五步提示词。