跳到正文

Agent 智能体

让模型自主规划、调用工具、完成多步任务的技术方向——从 Claude Code、Manus 到各家 Agent 框架与评测基准的全部动态。

当前仅显示精选新闻

最新精选

第 621–640 条 · 共 820 条
6月10日周三
  1. @op741872

    Anthropic 发布 Mythos 低配版 Claude Fable 5,向 API、Pro、Max、Team 及企业用户开放。它采用与 Mythos 5 相同的底层模型,API 输入每百万 Token 10 美元、输出 50 美元,比 Mythos Preview 便宜一半。Fable 加强了安全防护,涉及网络攻击、生化攻击或大规模能力蒸馏的请求会被拒绝,并回退到 4.8 版本。

    引用Claude (@claudeai)@claudeai

    Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. Video

    推荐理由:梳理了 Fable 5 与 Mythos 5 的开放范围、API 定价减半及安全回退机制,可据此评估其可用性。

  2. Claude Code GitHub Releases70

    Claude Code v2.1.170 发布:引入 Claude Fable 5 并修复会话转录保存问题

    Claude Code 发布 v2.1.170,宣布引入 Claude Fable 5,称其为面向通用场景开放、能力超过此前所有公开模型的 Mythos 级模型,升级该版本即可使用,并附 Anthropic 公告链接。该版本还修复了从 VS Code 集成终端或继承 Claude Code 环境变量的 shell 启动时,会话不保存转录且不出现在 --resume 中的问题。

    推荐理由:原文给出新模型的获取方式和一个会话转录保存的修复,Claude Code 用户可据此决定是否升级到该版本。

6月9日周二
  1. Hugging Face Blog66

    智能体如何串联两个 Hugging Face Space 搭建 3D 巴黎画廊

    一个编码智能体通过读取 Hugging Face Spaces 的 agents.md,串联图像生成 Space ideogram-ai/ideogram4 与单图 3D 重建 Space VAST-AI/TripoSplat,自动搭建出展示巴黎地标 3D Gaussian splat 的静态网站。

    推荐理由:原文记录了用 agents.md 让编码智能体串联图像与 3D 重建两个 Space 的完整流程,方法可直接复用。

  2. Claude Blog66

    Claude Managed Agents 新增定时部署与 Vault 环境变量存储

    Anthropic 为 Claude Managed Agents 新增定时部署和 Vault 两项功能,均已在 Claude Platform 进入公测。定时部署按 cron 调度运行智能体,每次触发启动新会话完成任务,无需自建调度器,部署上线后可随时暂停、恢复、归档或按需触发额外运行。

    推荐理由:官方给出定时部署与 Vault 两项能力的开放入口和客户实践,读者可据此判断智能体运维方式的变化。

  3. @berryxia67

    Kimi Code 开源编码智能体迎来重大升级,支持一行命令 CLI 安装、零配置快速启动,并可拖入视频作为编码上下文,实现参考视频转 LUT、长视频转短视频、屏幕录像转代码。插件系统可拉取股票价格、财报和学术论文,同时支持 ACP 协议接入 JetBrains 和 Zed,并提供自定义 hooks 扩展工作流,可搭配 Kimi K2.6 使用。

    引用Kimi Developers (@KimiDevs)@KimiDevs

    Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-recording-to-code, and more​ 🔹Plugins for stocks, financial reports, academic papers, with more coming​ 🔹Supports the ACP protocol, and works with JetBrains, Zed, and more​ 🔹Hooks for custom tools and workflows​ Try it with Kimi K2.6 👉 kimi.com/code Issues, plugin ideas, and PRs welcome! Community feedback helps shape what ships next.​🚀

    推荐理由:原文列出升级后的零配置安装、视频上下文与插件能力,读者可据此判断编码智能体的门槛变化。

  4. @berryxia70

    Kimi Work 在 macOS 和 Windows 上线,可在本地桌面并行运行最多 300 个 AI 智能体。它搭配 WebBridge 扩展在浏览器中完成搜索、滚动、点击和打字,并为财经场景原生调用 Yahoo Finance 与世界银行数据,无需复杂 API 配置。产品自带记忆系统记录用户偏好与历史决定,任务完成后直接输出 PPTX、Word、PDF、Excel 文件。

    引用Kimi.ai (@Kimi_Moonshot)@Kimi_Moonshot

    Meet Kimi Work - a local AI agent on your desktop that does the work for you. 🔹Native agent swarm: Up to 300 AI agents running in parallel on your local machine. 🔹Browser use: Paired with WebBridge extension, your agent will navigate websites in your browser: search, scroll, click, type and complete tasks. 🔹Built for Finance: Native global market data tool call from Yahoo Finance and World Bank - no complex API setup required. 🔹Memory system: Kimi Desktop keeps a running diary of your preferences, past decisions, and context to know you better. Available for macOS (Apple Silicon) and Windows. 🔗Try it now: kimi.com/products/kimi-work Video

    推荐理由:官方披露了本地代理规模、浏览器操作与财经数据原生调用等能力,读者可据此判断桌面端智能体的落地形态。

6月8日周一
  1. 量子位 · 微信公众号76

    OpenAI 高管称「Chat 已死」,ChatGPT 将改版为 Agent 超级应用

    OpenAI 高管提出「Chat 已死」,公司正推进 ChatGPT 诞生以来最大规模改版,目标是从聊天机器人变成个人 Agent 式的超级应用。改版由 Codex 承担 Agent 能力,其周活已超 500 万、非开发者用户占 20%,并新增可直接操作电脑、并行运行多个 Agent 等能力。动因是 ChatGPT 虽在 5 月突破 10 亿月活但多数用户免费,企业端支出正被 Claude 抢占。

    推荐理由:文章梳理了 OpenAI 把 ChatGPT 从聊天框转向 Agent 超级应用的路线,并用 Codex 与 Claude 的数据呈现其转型压力。

  2. Hugging Face Blog61

    OpenEnv 转为委员会共同治理,成为开源智能体 RL 的互操作层

    OpenEnv 宣布由 Meta-PyTorch、Reflection、Unsloth、Modal、Prime Intellect、Nvidia、Mercor、Fleet AI、Microsoft、Hugging Face 和 RadixArk 组成的委员会共同协调,项目地址迁移至 huggingface/OpenEnv。

    推荐理由:OpenEnv 由多家机构组成的委员会共同协调,并明确为 RL 环境的互操作层,读者可了解开源智能体训练的协作与协议设计。

6月5日周五
  1. 数字生命卡兹克 · 微信公众号79

    Anthropic 发布万字长文《当 AI 开始构建自己》,讨论递归自我改进

    Anthropic 研究院发布长文《当 AI 开始构建自己》,用公开基准和此前未披露的内部数据说明 AI 已在加速 AI 系统自身的开发。文中称 Anthropic 工程师平均每季度交付的代码量是 2021 至 2025 年间的 8 倍,Claude 在最开放任务上的成功率在 2026 年 5 月达到 76%,六个月内提高 50 个百分点。

    推荐理由:Anthropic 用内部数据展示 Claude 在写代码和做研究上的进展,读者可据此理解递归自我改进这一趋势的早期证据。

  2. @kimmonismus75

    Anthropic 在一篇博客中称 AI 进展快于其内部预期,模型能可靠独立完成的任务时长约每四个月翻倍,此前趋势为每七个月。博客提到 Claude 已编写 Anthropic 代码库中 80% 以上的合并代码,并描述 AI 自主设计后继模型的递归自我改进前景。X 用户 @kimmonismus 转述该文并认为,即便模型能力冻结在当前水平,社会仍会因现有模型扩散而出现重大变化。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:文中引用 Anthropic 对递归自我改进的判断,并列出任务时长翻倍周期与代码占比等数据,便于把握当前的 AI 进展速度。

6月4日周四
  1. @kimmonismus81

    NVIDIA 发布 Nemotron 3 Ultra,一款完全开源的 550B MoE 模型,激活参数 55B,权重、训练数据与完整配方全部公开。该模型采用混合 Mamba-Attention MoE 架构,NVIDIA 称其在长输出智能体任务上的吞吐量约为同类开源模型的 6 倍,同时保持相同准确率。

    引用NVIDIA AI (@NVIDIAAI)@NVIDIAAI

    Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. Video

    推荐理由:原文给出 550B 开源模型的权重、训练数据与完整配方,并说明其在长任务智能体上的吞吐表现,读者可据此判断开源前沿模型的可复现程度。

  2. @RyanLeeMiniMax66

    MiniMax 发布开放权重模型 MiniMax M3,官方称其是首个同时结合编码与智能体、100 万上下文、原生多模态三项前沿能力的开放权重模型。官方给出 59.0% SWE-Bench Pro、66.0% Terminal Bench 2.1 等成绩,通过 MiniMax Sparse Attention 将上下文扩展至 1M,并从 Step Zero 起原生多模态。作者补充 M3 目前位列 ArtificialAnalysis 第 8,因需并行开源 MSA 算子,权重将于下周晚些时候向所有人发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:官方给出三项前沿能力与基准成绩,并说明权重和 MSA 算子的开源时间,读者可据此判断开放节奏与可用范围。

  3. @arena80

    MiniMax M3 登入 Arena,在 Code Arena 前端编码榜排名第 7,得分 1531,与 GLM-5.1 接近。其定价为每 M token 输入 0.60 美元、输出 2.40 美元,Arena 称其在所属价位推动了性价比前沿。MiniMax 官方介绍称 M3 是首个同时具备三项前沿能力的开源权重模型,SWE-Bench Pro 59.0%、Terminal Bench 2.1 66.0%、MCP Atlas 74.2%,通过 Sparse Attention 将上下文扩展至 1M,原生多模态,权重与技术报告约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:MiniMax M3 在 Arena 前端编码榜位列第 7,与 GLM-5.1 接近,读者可据此比较同级价位模型的能力与定价。

  4. @runware65

    MiniMax 发布开源权重模型 M3,称其同时结合编码与智能体、1M 上下文和原生多模态三项能力。官方给出 SWE-Bench Pro 59.0%、Terminal Bench 2.1 66.0%、MCP Atlas 74.2% 等成绩,并称 MiniMax Sparse Attention 将上下文扩展至 1M。Runware 表示已可通过其 API 调用该模型,权重与技术报告约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:引用的发布信息列出 M3 在编码、智能体与 1M 上下文上的基准成绩,可供判断这款开源权重模型的能力组合。

  5. @kimmonismus81

    Google 发布 Gemma 4 12B 开源模型,采用 Apache 2.0 许可,可在 16GB 显存笔记本上本地运行,支持智能体推理、视觉与音频,作者称其质量接近 Google 的 26B 模型。

    引用Google (@Google)@Google

    Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run locally with just 16GB of VRAM. It’s open and accessible for everyone to use under a permissive Apache 2.0 license. This is all made possible by our new, unified architecture that removes separate multimodal encoders. Here’s how we did it 🧵

    推荐理由:它把视觉与音频编码器并入主干,让 12B 模型能在 16GB 显存本地运行,读者可据此判断端侧多模态的门槛变化。

  6. @googleaidevs75

    Google 发布 Gemma 4 12B,一款不含多模态编码器的统一模型,可直接在笔记本上运行。该模型定位在移动端 E4B 与更大的 26B MoE 模型之间,视觉和音频输入直接进入 LLM 主干,原生支持音频,采用 Apache 2.0 许可。官方称在 16GB VRAM 下可本地运行复杂多步工作流,性能接近 26B 模型。

    推荐理由:官方给出无编码器架构与 16GB VRAM 本地运行的定位,读者可据此判断端侧多模态模型的选型边界。