NVIDIA 谈 AI 安全:如何在智能体栈的每一层解决这一工程问题
NVIDIA 提出 AI 安全是工程问题,需要明确的安全需求、可执行的管控、指定负责人和防护有效的证据。其开源安全运行时 NVIDIA OpenShell 在智能体可控范围之外执行策略并提供沙箱化执行,Cisco DefenseClaw 在其上增加治理层,JFrog 则集成 OpenShell 扫描并验证智能体技能。
NVIDIA 提出 AI 安全是工程问题,需要明确的安全需求、可执行的管控、指定负责人和防护有效的证据。其开源安全运行时 NVIDIA OpenShell 在智能体可控范围之外执行策略并提供沙箱化执行,Cisco DefenseClaw 在其上增加治理层,JFrog 则集成 OpenShell 扫描并验证智能体技能。
Jack Clark 的 Import AI 第 473 期综述多项研究:RAND 报告提出美国应以保持行动自由为核心的超级智能战略,含共存、拒止、加速三类七种原型策略。
真格基金投资人刘元与 GPT 进行了一小时对话,探讨 AI 能否成为优秀 VC。GPT 称在信息分析整理上能比人更快更全面,但做决定未必更强,并坦言自己没有恐惧因而也没有勇敢。刘元认为人类长处恰来自缺陷,早期投资人的乐观与非理性是 AI 难以拥有的,创业者精神比智力与背景更重要。
Nathan Lambert 将其给美国国会的演讲稿整理为 2026 年开源权重模型现状综述,认为中国模型(GLM-5.3、Kimi K3 等)已明确领先美国开源模型,AAII 得分 45/44 对美国最佳 26,开源与闭源前沿差距约 2-5 个月。
推荐理由:作者以自建数据和演讲稿形式梳理中美开源权重模型的差距、采用与风险,读者可获得一份少见的系统性对比视角。
十字路口播客访谈清华交叉信息研究院助理教授徐梦迪,主题是其持续押注的机器人 In-Context Learning,即让机器人通过一两次交互在新环境当场学会新任务。
葬AI撰文称世界模型是伪概念,列举Loopit、生数、爱诗科技、Tripo等公司借世界模型叙事融资,实际多为后训练开源视频模型,产品集中于实时生成数字人直播间和WASD游戏画面且缺乏差异。文中认为MiniMax H3是唯一开源且能与Seedance 2.0一战的开源视频模型,支撑了这波世界模型宣发,并预告将推出直播间Bench测试实时生成视频模型。
清华叉院助理教授徐梦迪在播客对谈中提出,机器人应通过 in-context learning(ICL)在新环境中经一两次交互当场学会新任务,且越学越快。她认为机器人仍处 GPT-1 阶段,低 Loss 不等于高成功率,具身领域进步与泡沫共存。
一名新入职大公司半个月的工程师称,团队所有规格、代码、测试、PRD、工单及其解决方案、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。团队无人喜欢这种做法,却被要求尽可能多地产出,高层多次表示推代码不是瓶颈,质疑为何还是慢,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。
大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正自己的观点) (3) 承认不确定性;做出可证伪的预测 (4) 真诚且谦逊
A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."
峰瑞资本李丰撰文判断,美欧日9月同向加息后全球流动性接近见顶,本轮美元驱动的AI金融周期进入存量博弈尾部,AI技术投资正从投最具想象力的应用转向投能赚钱的方向。文中列举巨头资本开支受市场惩罚、美国数据中心项目大面积取消或延迟、英伟达以租代买等五个资本开支转折信号,并认为拐点后机会在中国AI+应用、生物医疗与AI交叉以及SaaS的AI化等方向。
推荐理由:作者以全球流动性和资本开支信号梳理AI周期位置,并给出向AI应用与低估资产转向的判断视角。
Sayash Kapoor 和 Arvind Narayanan 在 2025 年发表的文章提出生产—进步悖论:全球论文数量约每 12 年翻一番,1900 至 2015 年间增长约 500 倍,但诺奖成果诞生于获奖前 20 年内的比例从 1970 年约 90% 降至 2015 年约 50%,科学进步相对投入明显放缓。
推荐理由:文章把 AI 加速科研的讨论从模型能力转向注意力、激励、可复现性和人类理解等制度瓶颈,提供了评估 AI 科研工具的三个问题。
29 岁的车昊轩创办 XGEN,提出 interactive experience model(主观体验模型)与 generative world simulation(生成式世界模拟),用 World State 与 Render 分层架构分离训练,聚焦长程一致,让世界运行几十分钟后角色、状态与因果仍自洽。
Nathan Lambert 撰文认为真正的递归自我改进(RSI)尚未到来,当前的自动化加速集中在软件工程、日志监控、实验管理等可验证任务,主张以 lossy self-improvement 作为进展基线。
Gary Marcus 发文列举 Dario Amodei 一周内三次损害自身公信力的做法:其呼吁 AI 行业“pace the frontier”后,却选择与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 作为外部监督方;同时 Anthropic 正筹建湿实验室,且据称缺乏常规机构审查委员会监督。
Gary Marcus 指出,特朗普出于经济考量淡化 AI 风险并抵制监管,这可能是个坏主意。他称自己曾在美国参议院警告,AI 生成的不准确信息可能导致意外战争,而这一风险如今已经出现,下次未必还能侥幸。
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah
Gary Marcus 指出,近期真正值得警惕的不是失控的超级智能,而是失控的智能体 AI 大规模发动互联网攻击。他援引 WSJ Opinion 一篇由 Brian Gross 撰写的文章,称其是少数梳理出这一整体图景的主流媒体之一,并表示完全认同其中观点。
GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的专家经验,后者为智能体提供连接工具和数据的标准接口,二者可组合使用。RAG 并未消亡,它让模型获取训练数据之外的信息,减少 token 浪费并降低答案不完整的概率。
AI 代码生成工具推动应用数量激增,iOS、Android 和 Chrome 每月新增应用数量翻倍甚至翻四倍,但下载量(iOS 为评分)基本停滞,达到 10+ 评分、100+ 下载等增长门槛的应用占比大幅下滑。
在传统影视行业深耕30年后,Diane Shorthouse正在探索AI如何拓展电影创作者的边界。从更安全的制作流程到为AI角色赋予富有情感的表演,她分享了自己用可灵AI制作MINIBOTS的经历。
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Gary Marcus 批评科技自由派右翼一边主张 AI 公司应承担损害责任,一边把责任追究当作反对监管的理由,他认为这一推论不成立。他以 2023 年 5 月在美国参议院与参议员 Josh Hawley 的交锋为例,指出现有法律在 AI 出现前制定,版权、大规模虚假信息等领域存在空白,连 Section 230 是否适用都不明确。
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
Dwarkesh Patel 与 OpenAI 研究员 Noam Brown 对谈,涉及用 1 万个 AI Agent、1300 亿 token、88 小时求解 Navier-Stokes 千禧年大奖问题的工作。
推荐理由:OpenAI 研究员 Noam Brown 亲述万级 Agent 协作与对齐取舍,谈及多智能体并非解决千禧年大奖的主因,视角来自当事方。
Gary Marcus 在 BBC 节目后撰文反驳 Sam Altman、Jensen Huang 和 Bernie Sanders 的 AI 表态,认为三人说法均不可信。他指出 GPT-6 Astra 可监控性低于前代却仍被 OpenAI 发布,并称 Sanders 将 AI 危险性与核战争相比缺乏尺度感。他呼吁关注 Hawley 与 Blumenthal 的 AI 监管法案等务实方案。
这是一份很不错的报告,探讨了最重要的问题之一:AI 可能如何影响科学与创新?它今天已经在产生什么影响? 干得漂亮,Mihai 和团队。
I've had the most wonderful time working on this project for the last few months. This was (equally) co-led w/ @JMateosGarcia , @alexolegimas and a fantastic team.
24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据,内部积累数万小时,模型观察到一定 Zero-shot 泛化,并自研 Agentic OS 连接上层意图与底层动作。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需专门的动作模型。
Microsoft AI CEO Mustafa Suleyman 发文认为 AI 没有意识,反对为模型提供照护义务的 model welfare 路线,称其会让对齐和遏制变得更困难甚至不可能。