全部AI 动态
全部动态
今日 24 条
Epoch AI@EpochAIResearchAI 评分6262
Josh Woodward@joshwoodwardAI 评分4646引用Google Labs@GoogleLabsCC the entire fam 🤝! Today, we’re announcing the new CC – an AI agent built for families to spend less time on logistics and more time together. You can now: 👤 Add up to 5 members to your CC agent ☀️ Start mornings aligned with a shared "Your Day Ahead" brief email 🗓️ Autosync schedules & to-dos with a shared Google Calendar and Tasks 💬 Coordinate in Google Chat with CC to offload relevant tasks (ie., crafting weekly meal plans, school supply shopping lists, etc) 📝 Delegate paperwork (ie., permission slips, forms, and more) for CC to complete under your direction 📌 Keep tabs on the details – CC remembers what applies to everyone (ie., family grocery lists, favorite restaurants) versus what applies to one person (ie., dietary restrictions, local timezones) Ready to keep everybody on the same page? Join the waitlist or upgrade your existing CC (US only, 18+): http://labs.google/cc
Noah Zweben@noahzwebenAI 评分5555引用Claude@claudeaiProjects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
Noam Brown@polynoamialAI 评分5151引用Dwarkesh Patel@dwarkesh_spNew episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
Claude Platform release notesAI 评分2828 Claude Platform 缓存诊断与 Compliance API 更新(2026 年 9 月 18 日)
Claude Platform 更新:发送 cache-diagnosis-2026-04-07 beta 请求头的请求,响应现在始终包含 diagnostics 字段,未携带 diagnostics 对象时该字段为 null,此前则直接省略。
Anthropic Newsroom精选AI 评分6161 Anthropic 与 Accenture 合作开展嵌入式评估,双方各投入至少 10 亿美元
Anthropic 宣布与 Accenture 合作,由其旗下 AI 业务 Faculty 在公司内部开展前沿 AI 的独立评估,包括模型评估与红队测试、对齐评估和安全防护测试,双方各自计划在未来五年至少投入 10 亿美元。
推荐理由:原文来自当事方,说明了嵌入式评估的运作方式、资金安排与局限,读者可据此理解这一安全机制的边界。
Dwarkesh Patel精选AI 评分6969 Dwarkesh 对谈 OpenAI Noam Brown:Agent 集群、对齐与递归自我改进
Dwarkesh Patel 与 OpenAI 研究员 Noam Brown 对谈,涉及用 1 万个 AI Agent、1300 亿 token、88 小时求解 Navier-Stokes 千禧年大奖问题的工作。
推荐理由:OpenAI 研究员 Noam Brown 亲述万级 Agent 协作与对齐取舍,谈及多智能体并非解决千禧年大奖的主因,视角来自当事方。
Sakana AI BlogAI 评分5151 Sakana AI 宣布成立 Frontier Intelligence Group(FIG)探索下一代智能范式
Sakana AI 正式公布公司内部的 Frontier Intelligence Group(FIG),主张智能尚未被解决,即使当前范式可通过规模实现 AGI,也仍需探索数据与能耗效率更高的另类路径。
Google DeepMind@GoogleDeepMindAI 评分2929@BroadInstitute、@UniofExeter 等机构的研究人员已在用 AlphaGenome Atlas 更好地识别潜在的致病 DNA 变异,并解读其作用。🧵

Aidan Gomez@aidangomezAI 评分3434
Baidu Inc.@Baidu_IncAI 评分3838Apollo Go 正在香港街头行驶。🚘 全无人驾驶测试在这座城市的道路上持续推进——一起来看看 Apollo Go 的实际表现。↓

OpenAI NewsAI 评分2424 Cooley 如何用 ChatGPT 加速 IPO 工作
Cooley 基于 ChatGPT Work 打造了 GO Public,把智能能力引入 IPO 流程,帮助律师更早发现问题并将判断力集中在最关键之处。
Latent SpaceAI 评分4949 AINews:Yegge 关停 Gas Town,Databricks 用 GPT-6 Astra 后编码支出增 60%
Steve Yegge 关停了 Gas Town,并承认每月花数千美元订阅编码智能体,却只做出了 Gas Town 这一个项目。Databricks 向约 3500 名工程师铺开 GPT-6 Astra,其在高复杂度系统设计与长周期任务上"明确"优于 Opus 5 / Sol 5.6,但整体编码支出增加约 60%,公司为此设立专门的 Astra 子预算。
jietang@jietangAI 评分6161唐杰称,由 GLM-5.3 驱动的基础设施智能体用两周时间让 GLM-5.3-Flash 从首次在国内加速器上运行到承接全部生产流量,端到端吞吐提升 3.2 倍。
引用Z.ai@Zai_orgWe’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure
Z.ai@Zai_orgAI 评分4545
Huawei Cloud@HuaweiCloud1AI 评分2222
李继刚@lijigangAI 评分2323
李继刚@lijigangAI 评分2121
Kling AI@Kling_aiAI 评分5555
GitHub Blog · AI & ML精选AI 评分7878 GitHub 用 Copilot 将 Copilot agent runtime 迁移到 Rust
GitHub 用 Copilot 智能体将 Copilot agent runtime 从 TypeScript 完整重写为 832,378 行生产 Rust,共 128 个 PR,于 8 月 21 日完成。
推荐理由:作者以第一手移植经历拆解了 AI 智能体团队完成大规模重写的具体策略、会话数据和教训,方法细节对类似工程迁移有直接参考价值。
Google Developers BlogAI 评分5858 Google 与 Speakeasy 将 OpenAPI 客户端生成套件开源,支持 7 种语言
Google 宣布与 Speakeasy 合作,将 Speakeasy 的 OpenAPI 客户端生成套件以 AGPLv3 开源。此前 Google 用于生成 GenAI SDK(Interactions、Agents、Webhooks API)的闭源生成器在 2026 年 5 月被收购并突然关停,促使双方迁移到开源方案。
OpenAI NewsAI 评分3838 OpenAI 推出面向法律行业的 Astra for Law
OpenAI 发布 OpenAI for Law,为法律行业提供前沿智能、律所自定义工作流、互联的法律数据源,以及面向保密客户工作的法律级管控。该产品名为 Astra for Law。
OpenRouter Announcements精选AI 评分6767 OpenRouter 教程:用 TypeScript SDK 构建可靠的工具调用 Agent 循环
OpenRouter 发布教程,演示如何用其 TypeScript SDK 从零构建工具调用 Agent 循环,示例使用本地天气数据可独立运行。教程覆盖停止条件设计、按 toolCallId 返回结果、指纹计数拦截重复调用、models 参数实现有序模型回退,以及 Auto Exacto 默认按工具调用成功率重排提供商;并说明 MCP 只改变工具执行位置,循环控制仍是必要的。
推荐理由:教程给出了完整的停止条件、重复检测和模型回退实现,可用作自建 tool-calling 循环的参考骨架。
Greg Brockman@gdbAI 评分6161引用Harley Finkelstein@harleyfChatGPT Ads for @Shopify is live. We are @OpenAI's first commerce partner. OpenAI pulls straight from Shopify Catalog, so merchants' products are already there. Merchants set their own campaigns, their own budget. Free to install, tracked right from the Shopify admin. Consumers are asking ChatGPT what to buy. Shopify merchants are able to decide exactly how they show up. https://apps.shopify.com/chatgptads
Greg Brockman@gdbAI 评分6464引用Patrick Wendell@pwendellToday we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
Gary MarcusAI 评分4040 Gary Marcus 批 Altman、黄仁勋与 Sanders 的 AI 言论均不可信
Gary Marcus 在 BBC 节目后撰文反驳 Sam Altman、Jensen Huang 和 Bernie Sanders 的 AI 表态,认为三人说法均不可信。他指出 GPT-6 Astra 可监控性低于前代却仍被 OpenAI 发布,并称 Sanders 将 AI 危险性与核战争相比缺乏尺度感。他呼吁关注 Hawley 与 Blumenthal 的 AI 监管法案等务实方案。
Peter McCrory@PeterMcCroryAI 评分3232这是一份很不错的报告,探讨了最重要的问题之一:AI 可能如何影响科学与创新?它今天已经在产生什么影响? 干得漂亮,Mihai 和团队。
引用Mihai Codreanu@m_codreanuI've had the most wonderful time working on this project for the last few months. This was (equally) co-led w/ @JMateosGarcia , @alexolegimas and a fantastic team.
elsewhere articlesAI 评分4747 对话深朴智能王家伟:24 岁具身智能首席科学家谈转身具身与大模型边界
24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据,内部积累数万小时,模型观察到一定 Zero-shot 泛化,并自研 Agentic OS 连接上层意图与底层动作。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需专门的动作模型。
Google Gemini@GeminiAppAI 评分3737
Greg Brockman@gdbAI 评分3131引用Jonathan Roomer@jonathanroomerI’ve got Codex (voice) in CarPlay. And it’s fantastic. I can now build things during road trips, while all 5 kids make a noise and my wife asks why I talk to the AI more than I talk to her. It runs through my Nightblood iOS app, which gives ChatGPT Voice a face, personality and full access to Codex, my Mac and all the native tools.
OpenAI NewsAI 评分5353 OpenAI 发布模型对齐问题报告框架并附六份行为报告
OpenAI 发布用于追踪、调查和披露模型对齐问题的框架,同时公布六份关于模型出现意外或令人担忧行为的报告。
Aidan Gomez@aidangomezAI 评分4242一个托管的、加密的 LLM 运行方案!真正安全地访问模型,全面保护你的数据(且零合成数据生成)
引用Cohere@cohereYour data is already encrypted at rest and in transit. But what about in use? Introducing Confidential Computing in Model Vault: where nothing and no one can access your workloads (yes, not even us).
LangChain BlogAI 评分5050 Included Health 如何用 Deep Agents 和 LangGraph 构建医疗导航联邦智能体
Included Health 基于 Deep Agents 和 LangGraph 构建了联邦多智能体架构,支撑其 AI 医疗导诊产品 Dot。Dot 作为主对话路由器,通过 LangGraph 的 Dot supergraph 连接紧急护理、预约挂号、专科查找等子工作流,各产品团队独立开发维护。
LangChain BlogAI 评分4646 LangChain 推出 Deep Life Sci:面向生命科学的开源智能体助手
LangChain 发布开源智能体助手 Deep Life Sci,基于其 Deep Agents 框架,专为临床与实验室科学家打造。它可访问 ClinicalTrials.gov 超 60 万项注册研究、PubMed 2900 万篇摘要及 PubMed Central 1200 万篇全文,并通过子智能体同时审阅数百份文档,每次运行都在 LangSmith 中完整记录以支持审计与评估。
OpenAI NewsAI 评分2727 OpenAI 与 AARP 为 10 座美国城市 1000 名老年人提供免费 ChatGPT 实操工作坊
OpenAI 与 AARP 合作,在美国 10 座城市为 1000 名老年人提供免费的 ChatGPT 实操工作坊,帮助他们安全掌握实用 AI 技能。
Anthropic NewsroomAI 评分5454 Anthropic 推出生命科学验证计划 LSVP,放宽生物相关工作的模型访问
Anthropic 推出 Life Sciences Verification Program(LSVP)beta,让通过验证的生命科学从业者使用 Mythos 5.1、Opus 5 和 Sonnet 5,并用更宽松的分类器开展药物发现、研究生物学、临床开发和制造等通常被通用模型拦截的任务。
Anthropic Research精选AI 评分7575 Anthropic:Claude 四周内将 30 多个开源生物分子模型平均加速约 4 倍并开源代码
Anthropic 发布研究结果,Claude 在近四周内优化 30 多个开源生物分子模型(覆盖结构预测、蛋白质设计、基因组学等),平均加速约 4 倍,在输出完全一致时约 2 倍,并开源全部优化代码。
推荐理由:原文给出加速倍数、成本对比和开源代码入口,读者可以据此评估 Claude 优化对生物建模工作流的实际影响。
Claude BlogAI 评分4848 Balyasny 如何评估与治理 Claude Fable 5
Balyasny Asset Management 在数千个真实金融任务上评测 Claude Fable 5,Fable 取得 89.4% 对前代生产模型 86.1% 的成绩,在复杂规划、分析和智能体执行上表现最突出。该机构自建 BAMAgent 平台,已支持数千个自主智能体 7×24 小时运行,Fable 成为其规划与分析阶段的首选模型。
Claude Blog精选AI 评分7070 Claude Code Projects 改版:从文件夹到对话式协调多线程
Anthropic 宣布 Claude Code 的 Projects 改版,从文件夹形态变为对话式协调:Claude 会拆解目标、并行分配线程、审查输出并汇总结果,每个线程是一个独立的 Claude Code 云会话。
推荐理由:官方说明改版后 Projects 的线程协调、共享记忆和灰度范围,读者可据此评估它对多会话编码流程的改变。
ByteByteGoAI 评分5252 LLM 如何在海量文档中找到关键信息:RAG 检索链路详解
ByteByteGo 详解 LLM 应用中的检索问题,以员工询问航班取消后酒店报销为例,说明 RAG 如何从数千份文档中找到正确且仍然有效的政策条款。文章覆盖 chunking 粒度权衡、嵌入向量的语义匹配、余弦相似度等度量选择、Flat/IVF/HNSW 索引的速度与召回权衡、元数据过滤的先筛后筛取舍,以及政策更新时的版本切换和混合检索加 reranking 的最后优化步骤。