跳到正文

全部动态

今日 24 条
9月18日周五
  1. Josh Woodward46

    有家庭吗?我们家里就爱用这个。 @GoogleLabs: 全家一起 CC 🤝! 今天,我们宣布全新的 CC——一款为家庭打造的 AI 智能体,让大家少花时间处理琐事,多花时间相处。 你现在可以: 👤 最多添加 5 名成员到你的 CC 智能体 ☀️ 以共享的“Your Day Ahead”简报邮件开启每一天 🗓️ 通过共享的 Google Calendar 和 Tasks 自动同步日程与待办 💬 在 Google Chat 中与 CC 协作,分派相关任务(例如制定每周餐食计划、学校用品购物清单等) 📝 委派文书工作(例如许可单、表格等),由 CC 在你的指导下完成 📌 掌握细节——CC 会记住哪些适用于所有人(例如家庭购物清单、常去的餐厅),哪些只适用于某个人(例如饮食限制、本地时区) 准备好让全家人步调一致了吗?加入等待名单或升级你现有的 CC(仅限美国,18 岁以上):http://labs.google/cc

    引用Google Labs@GoogleLabs

    CC the entire fam 🤝! Today, we’re announcing the new CC – an AI agent built for families to spend less time on logistics and more time together. You can now: 👤 Add up to 5 members to your CC agent ☀️ Start mornings aligned with a shared "Your Day Ahead" brief email 🗓️ Autosync schedules & to-dos with a shared Google Calendar and Tasks 💬 Coordinate in Google Chat with CC to offload relevant tasks (ie., crafting weekly meal plans, school supply shopping lists, etc) 📝 Delegate paperwork (ie., permission slips, forms, and more) for CC to complete under your direction 📌 Keep tabs on the details – CC remembers what applies to everyone (ie., family grocery lists, favorite restaurants) versus what applies to one person (ie., dietary restrictions, local timezones) Ready to keep everybody on the same page? Join the waitlist or upgrade your existing CC (US only, 18+): http://labs.google/cc

  2. Noah Zweben55

    Claude 推出 Projects 功能,可在单一对话中管理项目,从 Claude Code 开始,由 Claude 调度多个并行线程,用户合上电脑后任务继续运行。该功能今日起对部分 Pro 和 Max 用户的云会话开放 Beta,将逐步向所有 Claude 用户推出。作者评论称通过单一协调者收集上下文、例行流程和输出,对管理复杂任务非常有价值。

    引用Claude@claudeai

    Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.

  3. Noam Brown51

    OpenAI 的 Noam Brown 在 Dwarkesh 播客中深谈多智能体、Navier-Stokes 与当前数学进展对自动化 AI 研究和递归自我改进的启示。讨论还涵盖如何在启动 RSI 前判断模型是否真正对齐,以及思维链退化、内外部模型差距等话题,并感谢 OpenAI 团队在多智能体方面的工作。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?

  4. Anthropic Newsroom61

    Anthropic 与 Accenture 合作开展嵌入式评估,双方各投入至少 10 亿美元

    Anthropic 宣布与 Accenture 合作,由其旗下 AI 业务 Faculty 在公司内部开展前沿 AI 的独立评估,包括模型评估与红队测试、对齐评估和安全防护测试,双方各自计划在未来五年至少投入 10 亿美元。

    推荐理由:原文来自当事方,说明了嵌入式评估的运作方式、资金安排与局限,读者可据此理解这一安全机制的边界。

9月17日周四
  1. jietang61

    唐杰称,由 GLM-5.3 驱动的基础设施智能体用两周时间让 GLM-5.3-Flash 从首次在国内加速器上运行到承接全部生产流量,端到端吞吐提升 3.2 倍。

    引用Z.ai@Zai_org

    We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure

  2. GitHub Blog · AI & ML78

    GitHub 用 Copilot 将 Copilot agent runtime 迁移到 Rust

    GitHub 用 Copilot 智能体将 Copilot agent runtime 从 TypeScript 完整重写为 832,378 行生产 Rust,共 128 个 PR,于 8 月 21 日完成。

    推荐理由:作者以第一手移植经历拆解了 AI 智能体团队完成大规模重写的具体策略、会话数据和教训,方法细节对类似工程迁移有直接参考价值。

  3. OpenRouter Announcements67

    OpenRouter 教程:用 TypeScript SDK 构建可靠的工具调用 Agent 循环

    OpenRouter 发布教程,演示如何用其 TypeScript SDK 从零构建工具调用 Agent 循环,示例使用本地天气数据可独立运行。教程覆盖停止条件设计、按 toolCallId 返回结果、指纹计数拦截重复调用、models 参数实现有序模型回退,以及 Auto Exacto 默认按工具调用成功率重排提供商;并说明 MCP 只改变工具执行位置,循环控制仍是必要的。

    推荐理由:教程给出了完整的停止条件、重复检测和模型回退实现,可用作自建 tool-calling 循环的参考骨架。

  4. Greg Brockman61

    ChatGPT Ads for Shopify 上线,Shopify 是 OpenAI 的首个商务合作伙伴,Greg Brockman 转发宣布该消息。OpenAI 直接从 Shopify Catalog 拉取商品数据,商家可自行设置广告活动和预算,免费安装并可在 Shopify 后台追踪效果。

    引用Harley Finkelstein@harleyf

    ChatGPT Ads for @Shopify is live. We are @OpenAI's first commerce partner. OpenAI pulls straight from Shopify Catalog, so merchants' products are already there. Merchants set their own campaigns, their own budget. Free to install, tracked right from the Shopify admin. Consumers are asking ChatGPT what to buy. Shopify merchants are able to decide exactly how they show up. https://apps.shopify.com/chatgptads

  5. Greg Brockman64

    Databricks 将 Astra 推广到全部约 3500 名工程师。其内部试点约 200 人的数据显示,Astra 在复杂任务上明显优于此前最强的 Opus 5 和 Sol 5.6,使用 Astra 的工程师整体编码支出增加约 60%,但在中低复杂度任务上相比既有模型提升不明显。团队通过 Unity Gateway 做分群实验,并为 Astra 单独设预算,鼓励工程师在复杂任务上选用 Astra、日常任务用更低成本模型;因数据保留政策尚未广泛部署 Fable,暂无 Astra 与 Fable 的可靠对比。

    引用Patrick Wendell@pwendell

    Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.

  6. elsewhere articles47

    对话深朴智能王家伟:24 岁具身智能首席科学家谈转身具身与大模型边界

    24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据,内部积累数万小时,模型观察到一定 Zero-shot 泛化,并自研 Agentic OS 连接上层意图与底层动作。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需专门的动作模型。

  7. Greg Brockman31

    Codex voice for road trips:

    引用Jonathan Roomer@jonathanroomer

    I’ve got Codex (voice) in CarPlay. And it’s fantastic. I can now build things during road trips, while all 5 kids make a noise and my wife asks why I talk to the AI more than I talk to her. It runs through my Nightblood iOS app, which gives ChatGPT Voice a face, personality and full access to Codex, my Mac and all the native tools.

  8. LangChain Blog46

    LangChain 推出 Deep Life Sci:面向生命科学的开源智能体助手

    LangChain 发布开源智能体助手 Deep Life Sci,基于其 Deep Agents 框架,专为临床与实验室科学家打造。它可访问 ClinicalTrials.gov 超 60 万项注册研究、PubMed 2900 万篇摘要及 PubMed Central 1200 万篇全文,并通过子智能体同时审阅数百份文档,每次运行都在 LangSmith 中完整记录以支持审计与评估。

  9. Anthropic Research75

    Anthropic:Claude 四周内将 30 多个开源生物分子模型平均加速约 4 倍并开源代码

    Anthropic 发布研究结果,Claude 在近四周内优化 30 多个开源生物分子模型(覆盖结构预测、蛋白质设计、基因组学等),平均加速约 4 倍,在输出完全一致时约 2 倍,并开源全部优化代码。

    推荐理由:原文给出加速倍数、成本对比和开源代码入口,读者可以据此评估 Claude 优化对生物建模工作流的实际影响。

  10. Claude Blog48

    Balyasny 如何评估与治理 Claude Fable 5

    Balyasny Asset Management 在数千个真实金融任务上评测 Claude Fable 5,Fable 取得 89.4% 对前代生产模型 86.1% 的成绩,在复杂规划、分析和智能体执行上表现最突出。该机构自建 BAMAgent 平台,已支持数千个自主智能体 7×24 小时运行,Fable 成为其规划与分析阶段的首选模型。

  11. Claude Blog70

    Claude Code Projects 改版:从文件夹到对话式协调多线程

    Anthropic 宣布 Claude Code 的 Projects 改版,从文件夹形态变为对话式协调:Claude 会拆解目标、并行分配线程、审查输出并汇总结果,每个线程是一个独立的 Claude Code 云会话。

    推荐理由:官方说明改版后 Projects 的线程协调、共享记忆和灰度范围,读者可据此评估它对多会话编码流程的改变。

9月16日周三
  1. ByteByteGo52

    LLM 如何在海量文档中找到关键信息:RAG 检索链路详解

    ByteByteGo 详解 LLM 应用中的检索问题,以员工询问航班取消后酒店报销为例,说明 RAG 如何从数千份文档中找到正确且仍然有效的政策条款。文章覆盖 chunking 粒度权衡、嵌入向量的语义匹配、余弦相似度等度量选择、Flat/IVF/HNSW 索引的速度与召回权衡、元数据过滤的先筛后筛取舍,以及政策更新时的版本切换和混合检索加 reranking 的最后优化步骤。