跳到正文

Agent 智能体

让模型自主规划、调用工具、完成多步任务的技术方向——从 Claude Code、Manus 到各家 Agent 框架与评测基准的全部动态。

当前仅显示精选新闻

最新精选

第 681–700 条 · 共 812 条
5月28日周四
  1. Mistral AI71

    Mistral 发布统一智能体 Vibe,Le Chat 升级为 Work 与 Code 双模式

    Mistral 将 Le Chat 升级为统一智能体 Vibe,一个账号覆盖工作与编码场景,已上线 chat.mistral.ai。Work Mode 面向长时多步任务,支持企业知识搜索、数据分析、文档起草、定时任务和可复用技能;Code Mode 提供远程编码会话并推出 VS Code 扩展,CLI 新增技能命令、会话权限控制和 /teleport 迁移。

    推荐理由:原文是 Mistral 官方发布,完整列出 Work 与 Code 两种模式的能力、集成入口和各档定价,读者可据此评估是否替换现有工作流。

  2. Linear Now63

    Linear 发布 Diffs,主打快速代码评审

    Linear 发布代码评审产品 Diffs,面向所有套餐开放,目标是让代码评审在不牺牲严谨性的前提下保持快速。评审在 Linear 内打开并关联对应的 issue 和项目,优先级可见;Guided reviews 把大 diff 按工作的推理顺序拆成章节,先展示核心改动,结构化 diff 高亮会剥离格式改动,让 2000 行的 PR 在一次阅读中变得可读。

    推荐理由:Linear 官方解释了 Diffs 针对的评审痛点与具体设计,读者可以对照自己团队被 agent 大量 PR 拖慢的评审流程判断是否适用。

  3. @swyx74

    swyx 评论 Cognition 的融资,称其已是全球最大的独立智能体实验室,并建议读者用图表中 200% 的用量指标推算其销售增长。他列举了模型多样性、云开发基础设施、代码审查与安全、GTM 等优势,认为这是 Peter Thiel 最大的 AI 押注。

    引用Cognition (@cognition)@cognition

    1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.

    推荐理由:作者从模型多样性、GTM 与企业级验证等角度拆解 Cognition 的护城河,可结合其融资数据理解这轮投资逻辑。

  4. @swyx76

    swyx 评论 Cognition 完成超 10 亿美元融资、估值 260 亿美元,称其已是全球最大的独立智能体实验室。据其引用的 Cognition 公告,本轮由 Lux Capital、General Catalyst 和 8VC 领投,企业使用量自今年初增长超 10 倍,run-rate 收入达 4.92 亿美元。

    引用Cognition (@cognition)@cognition

    1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.

    推荐理由:swyx 给出 Cognition 融资后的投资逻辑,解释编码智能体为何被视为多重趋势叠加的产物。

  5. Hugging Face Blog68

    Artificial Analysis 与 IBM 发布 ITBench-AA 基准,前沿模型在智能体企业 IT 任务上得分不足 50%

    Artificial Analysis 与 IBM 软件创新实验室发布 ITBench-AA,这是首个面向智能体企业 IT 任务的基准系列,首批聚焦 SRE 场景,前沿模型得分均低于 50%。

    推荐理由:前沿模型在这套企业IT智能体基准上全部低于50%,并给出开源权重模型的每任务成本对照,可作为能力与成本参照。

5月27日周三
  1. Claude Blog76

    Anthropic 分享用 LLM 保障源码安全的六步流程

    Anthropic 发布指南,介绍如何用 Claude Opus 通过威胁建模、沙箱、发现、验证、分诊、修补六个步骤查找并修复源码漏洞。文中称截至 2026 年 5 月 22 日其开源软件扫描已披露 1596 个漏洞,其中 97 个已修补,瓶颈已从发现转向验证、分诊与修补。

    推荐理由:文章把漏洞发现到修补拆成六步方法并附可复用 skill 与仓库,瓶颈已转向验证与修补。

  2. MiniMax Blog62

    MiniMax Agent 升级为 Mavis 并推出 Agent Team 多智能体协作

    MiniMax 发布 Agent 整体升级并更名 Mavis,推出 Agent Teams,支持桌面端并行运行多角色 Agent 协作完成长任务,并将 TokenPlan 与 Agent Plan 合并为统一订阅,Agent 与 API 共享 Credits。

    推荐理由:官方详解 Agent Team 的角色分工、状态机设计和上下文成本等工程取舍,读者可以据此理解多智能体产品化的关键问题。

5月26日周二
  1. @berryxia70

    xAI 宣布 Grok Build 已面向全体 SuperGrok 及 X Premium+ 用户开放 Beta 版本。用户可使用计划模式(Plan Mode),通过 Imagine 生成图像与视频,并借助命令行工具(CLI)搭建自动化程序或编排器,入口为 x.ai/cli。

    引用xAI (@xai)@xai

    Grok Build is now available in Beta for all SuperGrok and X Premium+ users. Use Plan Mode, create images and videos with Imagine, and build automations or orchestrators with the CLI. Visit x.ai/cli to get started. Video

    推荐理由:xAI 把 Grok Build 开放给 SuperGrok 与 X Premium+ 用户,可了解其计划模式与 CLI 编排能力。

  2. @vista868

    作者指出只安装 Skill 还不够,为更好触发和应用,需要把 Skill 写入 Agent.md,并给出安装更新 Waza 的提示词,要求以后各种开发设计优先使用这套 skill。转引内容显示 Waza 支持 Claude Code、Codex、Cursor 和 Pi 作为 agent 运行时,包含 8 个 skill,无框架、无遥测。

    引用Tw93 (@HiTw93)@HiTw93

    🥷 Engineering habits you already know, turned into skills AI agents can run. Waza absorbed a mass of real project lessons recently. Now just as sharp for Mac native apps, CLI tools, and Rust as it is for web. Supports Claude Code, Codex, Cursor, and Pi as agent runtimes. Reviews your CLI like a shipped product. Debugs "works in source tree, breaks after install." Sweeps sibling instances after every fix. Blocks "fixed" until runtime evidence is verified. 25 anti-patterns, destructive command safety, treats fetched content as untrusted data. 8 skills, no framework, no telemetry. Your superpower prompt collection can be uninstalled. Too heavy. github.com/tw93/Waza

    推荐理由:原文给出把 Skill 写入 Agent.md 的提示词写法,读者可据此改善技能触发与应用效果。

5月25日周一
  1. @kimmonismus65

    代号 TrapDoor 的供应链攻击同时针对 npm、PyPI 和 Crates.io,投递 34 个恶意包,目标是加密货币、AI 和安全开发者,用于窃取钱包、SSH 密钥和云凭证。

    引用Socket (@SocketSecurity)@SocketSecurity

    More analysis, package details, IOCs, and GitHub-related activity here, including attacker-hosted payload/config infrastructure and PRs attempting to add .cursorrules / CLAUDE.md files to popular AI and developer projects: socket.dev/blog/trapdoor-cry…

    推荐理由:该攻击把 CLAUDE.md 与 .cursorrules 配置文件当作新入口,让安全从业者了解 AI 编程助手被利用的攻击路径。

5月24日周日
5月23日周六
  1. AI as Normal Technology65

    AI as Normal Technology 剖析 Google 智能体 916 美元造操作系统宣称的漏洞

    Sayash Kapoor 等人分析 Google 在开发者大会上发布 Gemini 3.5 Flash 和 Antigravity 2.0 时宣称的智能体团队以单一提示词、约 916.92 美元 API 费用和 2.6B tokens 造出操作系统的实验。

    推荐理由:文章逐条拆解 Google 智能体造操作系统的宣称,指出单一提示词等说法缺乏关键细节,并探讨开放世界评测需要的方法规范。

5月22日周五
  1. Mistral AI69

    Mistral 发布 Mistral Medium 3.5 并在 Vibe 和 Le Chat 推出云端远程智能体

    Mistral 发布 Mistral Medium 3.5,一个 128B 稠密开源权重模型(modified MIT 许可),256k 上下文窗口,SWE-Bench Verified 得分 77.6%,τ³-Telecom 得分 91.4,自托管最少只需四块 GPU,API 定价每百万输入 token $1.5、输出 token $7.5。

    推荐理由:官方同步发布模型与云端异步智能体,给出基准分数、定价和开源权重,可对照评估其编码与长程任务能力。

  2. @swyx69

    OpenAI 的 Codex goal mode 从实验功能转为正式功能,用户可在 Codex app、IDE 扩展或 CLI 中设定一个具体里程碑,Codex 会持续工作直到达成,跨越数小时甚至数天。过程中可以查看并引导进度,也可以暂停 Codex,还能开启侧边对话了解已完成的工作而不打断主任务。swyx 就此评论说,现在可以在目标执行中途暂停和调整。

    引用OpenAI Developers (@OpenAIDevs)@OpenAIDevs

    🥅 /goal has graduated from an experiment—for tasks big and small, Codex gets your work done. Use goal mode in the Codex app, IDE Extension, or CLI to give Codex a specific milestone, and it will keep working until it gets there, even across hours or days. You can check in and steer, and even pause Codex along the way. Pro tip: start side chats to understand the work that has been done so far without having to interrupt the main task. developers.openai.com/codex/… Video

    推荐理由:原文说明 Codex 的 goal mode 已从实验转为正式功能,读者可了解它如何支撑跨小时到数天的长任务。