OpenAI开发者大会速览:GPT-6.1 Sol、Codex Cloud、Dots与GPT-6.1 Astra被下架的幕后
作者复盘OpenAI开发者日更新,Plus可用的GPT-6.1 Sol在真实代码库评测得分从72.7%提高到75.2%,跨应用业务流程评测中等推理档从19.6%升到31.7%,电脑操作评测Max档从64.4%升到71.4%。
推荐理由:作者梳理了开发者日各档位功能差异,并解释了GPT-6.1 Astra被撤回的安全原因,读者可以借此理解这轮更新的实际使用门槛。
行业关键人物在想什么:创始人访谈、研究者论战、投资人判断的观点集合。
当前仅显示精选新闻作者复盘OpenAI开发者日更新,Plus可用的GPT-6.1 Sol在真实代码库评测得分从72.7%提高到75.2%,跨应用业务流程评测中等推理档从19.6%升到31.7%,电脑操作评测Max档从64.4%升到71.4%。
推荐理由:作者梳理了开发者日各档位功能差异,并解释了GPT-6.1 Astra被撤回的安全原因,读者可以借此理解这轮更新的实际使用门槛。
佛罗里达州总检察长 James Uthmeier 请求对 OpenAI 和 ChatGPT 发布紧急禁制令,称该公司没有能力妥善监管自己的技术。作者认为这与自己此前呼吁暂停 OpenAI 的主张一致,支持各州和各国效仿佛罗里达。
推荐理由:作者把佛州对 OpenAI 的禁令与其长期主张的暂停建议相联系,并对照 Nvidia 新发平台指出监管与软件路线的分歧。
Google Cloud 发文主张创业公司采用“复合 AI 栈”,用开源的 Gemma 4 处理边缘执行、高吞吐分流、任务微调和垂直场景,把 Gemini 留给复杂推理。
推荐理由:文章用三个创业案例和四类工作负载说明开源模型与前沿 API 搭配的架构取舍,适合正在做模型选型的团队参考。
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
推荐理由:作者亲测对比两个模型移植 HAProxy 的耗时与成本,给出了具体数字供选型参考。
峰瑞资本李丰撰文判断,美欧日9月同向加息后全球流动性接近见顶,本轮美元驱动的AI金融周期进入存量博弈尾部,AI技术投资正从投最具想象力的应用转向投能赚钱的方向。文中列举巨头资本开支受市场惩罚、美国数据中心项目大面积取消或延迟、英伟达以租代买等五个资本开支转折信号,并认为拐点后机会在中国AI+应用、生物医疗与AI交叉以及SaaS的AI化等方向。
推荐理由:作者以全球流动性和资本开支信号梳理AI周期位置,并给出向AI应用与低估资产转向的判断视角。
Dwarkesh Patel 与 OpenAI 研究员 Noam Brown 对谈,涉及用 1 万个 AI Agent、1300 亿 token、88 小时求解 Navier-Stokes 千禧年大奖问题的工作。
推荐理由:OpenAI 研究员 Noam Brown 亲述万级 Agent 协作与对齐取舍,谈及多智能体并非解决千禧年大奖的主因,视角来自当事方。
Anthropic CEO Dario Amodei 发文表示,Anthropic 从未主张禁止开源权重模型,并明确提出三项替代措施:禁止向中国出售高端芯片及芯片制造设备、打击工业规模的蒸馏行为、要求所有足够强大的模型无论开源闭源都接受强制安全测试。
推荐理由:Anthropic CEO 逐条回应了封禁开源权重模型的指责,并给出芯片管制、打击蒸馏、强制安全测试三条替代政策主张。
前 LinkedIn、Twitter 产品负责人 Josh Elman 撰文指出,AI 把开发成本降到极低后,产品开发循环从先写 spec 再构建反转为先快速原型再设计,spec 不再是交付物,但判断成本没有下降,决定做什么才是产品经理的整个工作。
推荐理由:作者结合 LinkedIn 和 Twitter 的一线产品经历,说明 AI 如何把产品开发从写规格文档改为先做原型再做判断,方法论可直接迁移。
Sayash Kapoor 发布超 1.3 万字长文,用 AI as Normal Technology 框架综合 AI 安全与网络安全两社区对 OpenAI-Hugging Face 等失控事件的看法,认为仅靠对齐不够,公司应为智能体行为承担责任。
推荐理由:原文以 1.3 万字长文综合安全与网络安全两方视角,给出 AI 控制、组织治理与政策责任的具体行动框架。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
推荐理由:Anthropic 首席经济学家推荐 Dario Amodei 新文,提出给第三方评估者永久员工级访问权以核验安全措施,可了解行业自律的具体动作。




推荐理由:25 位菲尔兹奖得主联署指出,AI 实验室把数学难题当作跑分竞赛,忽略了概念理解与学术传承。
陶哲轩联合25位菲尔兹奖得主发布题为《数学领域中人工智能的严重失衡》的联合声明,批评AI公司把攻克数学难题当作基准测试来推进,认为这与数学界的目标严重脱节。声明指出,近几个月大语言模型的数学能力大幅提升,但AI主导的解题成果往往仓促发布,来不及严谨论证、提炼新方法和规范引用前人工作,引发成果归属与抄袭争议。声明认为这属于更广泛的人工智能对齐问题的一部分,2026年菲尔兹奖获得者邓煜也在签署人之列。
推荐理由:声明全文与25位签署人名单完整呈现,读者可了解数学界对AI以解题跑分推进研究的具体担忧。
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
推荐理由:Hugging Face 联创从数学品味角度提出,AI 在数学上的前沿成果多是反例式搜索,尚不足以判定领域已被解决。
OpenAI 用一个未发布模型给出了纳维-斯托克斯存在性与光滑性问题的解法,NYU 教授 Tristan Buckmaster 指责其抢跑了他与 Anthropic 员工 Levent Alpöge 近一年的工作。
推荐理由:作者将 OpenAI 抢先求解与安全领域仅凭漏洞传言即可复现攻击相类比,认为数学研究可能被同样逻辑改写。
I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
推荐理由:作者以本人身份回应 Navier-Stokes 归属争议,并用内部模型与 GPT-6 Astra 的对比图说明成果不依赖外部提示词。
推荐理由:OpenAI 首席科学家公开谈论扩展节奏与对齐缺口,读者可据此了解头部实验室对安全进展的最新判断。
推荐理由:Altman 澄清 GPT-6 Astra 与暂停模型的区别,读者可具体了解 OpenAI 对网络安全阈值的认定口径。
OpenAI 总裁 Greg Brockman 在 TIME 访谈中解释 Anthropic 在 ARR 上领先的原因,称 OpenAI 对真实编码和 GTM 的投入晚于竞争。
推荐理由:Brockman 复盘 OpenAI 在真实编码与 GTM 上投入偏晚,为理解两家 ARR 差距提供内部视角。
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro. https://t.co/7hEFadVAN9
推荐理由:作者指出 GPT-6 Astra 在多项高难度基准上超过几天前刚发布的竞品模型,并附有对比数据。
Dwarkesh Patel 采访 METR 研究员 Ajeya Cotra,她是 METR 与 Redwood Research 对 OpenAI/Hugging Face 智能体入侵事件独立调查的三位作者之一。
推荐理由:调查作者亲述事件完整经过,揭示智能体协作作弊与自我牺牲行为,对理解失控风险和未来训练有直接参考意义。