推荐理由:整理了 Anthropic Computer Use 在分辨率、请求顺序和 thinking 预算上的具体建议,可对照调整截图与参数配置。
Agent 智能体
全部主题让模型自主规划、调用工具、完成多步任务的技术方向——从 Claude Code、Manus 到各家 Agent 框架与评测基准的全部动态。
当前仅显示精选新闻最新精选
第 681–700 条 · 共 812 条@vista8@vista8精选AI 评分7070 Mistral AI精选AI 评分7171 Mistral 发布统一智能体 Vibe,Le Chat 升级为 Work 与 Code 双模式
Mistral 将 Le Chat 升级为统一智能体 Vibe,一个账号覆盖工作与编码场景,已上线 chat.mistral.ai。Work Mode 面向长时多步任务,支持企业知识搜索、数据分析、文档起草、定时任务和可复用技能;Code Mode 提供远程编码会话并推出 VS Code 扩展,CLI 新增技能命令、会话权限控制和 /teleport 迁移。
推荐理由:原文是 Mistral 官方发布,完整列出 Work 与 Code 两种模式的能力、集成入口和各档定价,读者可据此评估是否替换现有工作流。
Linear Now精选AI 评分6363 Linear 发布 Diffs,主打快速代码评审
Linear 发布代码评审产品 Diffs,面向所有套餐开放,目标是让代码评审在不牺牲严谨性的前提下保持快速。评审在 Linear 内打开并关联对应的 issue 和项目,优先级可见;Guided reviews 把大 diff 按工作的推理顺序拆成章节,先展示核心改动,结构化 diff 高亮会剥离格式改动,让 2000 行的 PR 在一次阅读中变得可读。
推荐理由:Linear 官方解释了 Diffs 针对的评审痛点与具体设计,读者可以对照自己团队被 agent 大量 PR 拖慢的评审流程判断是否适用。
Anthropic Newsroom精选AI 评分7575 Anthropic 发布 Claude Opus 4.8,并上线动态工作流与 effort 控制
Anthropic 发布 Claude Opus 4.8,在编码、智能体与推理等基准上较 Opus 4.7 有提升,常规使用价格与 Opus 4.7 相同,为每百万输入 token 5 美元、输出 25 美元。
推荐理由:官方给出与前代及 GPT-5.5、Gemini 3.1 Pro 的基准对比,读者可据此判断它在编码与智能体任务中的能力位置。
Claude Blog精选AI 评分8383 Claude Code 推出动态工作流,单会话可并行调度上百个子智能体
Anthropic 为 Claude Code 推出动态工作流,Claude 可动态编写编排脚本,在单次会话中并行运行数十到数百个子智能体,并在结果交付前自行校验。
推荐理由:官方给出了并行子智能体编排的能力边界与 Bun 重写中的实际用例,便于判断它能否承接超长任务。
@swyx@swyx精选AI 评分7474
引用Cognition (@cognition)@cognition1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.
推荐理由:作者从模型多样性、GTM 与企业级验证等角度拆解 Cognition 的护城河,可结合其融资数据理解这轮投资逻辑。
@swyx@swyx精选AI 评分7676
引用Cognition (@cognition)@cognition1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.
推荐理由:swyx 给出 Cognition 融资后的投资逻辑,解释编码智能体为何被视为多重趋势叠加的产物。
Hugging Face Blog精选AI 评分6868 Artificial Analysis 与 IBM 发布 ITBench-AA 基准,前沿模型在智能体企业 IT 任务上得分不足 50%
Artificial Analysis 与 IBM 软件创新实验室发布 ITBench-AA,这是首个面向智能体企业 IT 任务的基准系列,首批聚焦 SRE 场景,前沿模型得分均低于 50%。
推荐理由:前沿模型在这套企业IT智能体基准上全部低于50%,并给出开源权重模型的每任务成本对照,可作为能力与成本参照。
Claude Platform release notes精选AI 评分8383 Anthropic 发布 Claude Opus 4.8,默认支持 1M token 上下文窗口
Anthropic 发布 Claude Opus 4.8,为其最强模型,在 Claude API、Amazon Bedrock、Google Cloud 和 Microsoft Foundry 上默认支持 1M token 上下文窗口,最大输出 128k token。
推荐理由:原文完整列出 1M 上下文、采样参数限制和各端可用性等迁移要点,接入方可以直接对照判断升级路径。
Claude Blog精选AI 评分7676 Anthropic 分享用 LLM 保障源码安全的六步流程
Anthropic 发布指南,介绍如何用 Claude Opus 通过威胁建模、沙箱、发现、验证、分诊、修补六个步骤查找并修复源码漏洞。文中称截至 2026 年 5 月 22 日其开源软件扫描已披露 1596 个漏洞,其中 97 个已修补,瓶颈已从发现转向验证、分诊与修补。
推荐理由:文章把漏洞发现到修补拆成六步方法并附可复用 skill 与仓库,瓶颈已转向验证与修补。
MiniMax Blog精选AI 评分6262 MiniMax Agent 升级为 Mavis 并推出 Agent Team 多智能体协作
MiniMax 发布 Agent 整体升级并更名 Mavis,推出 Agent Teams,支持桌面端并行运行多角色 Agent 协作完成长任务,并将 TokenPlan 与 Agent Plan 合并为统一订阅,Agent 与 API 共享 Credits。
推荐理由:官方详解 Agent Team 的角色分工、状态机设计和上下文成本等工程取舍,读者可以据此理解多智能体产品化的关键问题。
@berryxia@berryxia精选AI 评分7070 引用xAI (@xai)@xaiGrok Build is now available in Beta for all SuperGrok and X Premium+ users. Use Plan Mode, create images and videos with Imagine, and build automations or orchestrators with the CLI. Visit x.ai/cli to get started. Video
推荐理由:xAI 把 Grok Build 开放给 SuperGrok 与 X Premium+ 用户,可了解其计划模式与 CLI 编排能力。
@vista8@vista8精选AI 评分6868
引用Tw93 (@HiTw93)@HiTw93🥷 Engineering habits you already know, turned into skills AI agents can run. Waza absorbed a mass of real project lessons recently. Now just as sharp for Mac native apps, CLI tools, and Rust as it is for web. Supports Claude Code, Codex, Cursor, and Pi as agent runtimes. Reviews your CLI like a shipped product. Debugs "works in source tree, breaks after install." Sweeps sibling instances after every fix. Blocks "fixed" until runtime evidence is verified. 25 anti-patterns, destructive command safety, treats fetched content as untrusted data. 8 skills, no framework, no telemetry. Your superpower prompt collection can be uninstalled. Too heavy. github.com/tw93/Waza
推荐理由:原文给出把 Skill 写入 Agent.md 的提示词写法,读者可据此改善技能触发与应用效果。
Claude Blog精选AI 评分6161 Claude Managed Agents 新增自托管沙箱与 MCP 隧道
Claude Managed Agents 现在可在用户自行掌控的沙箱中运行,并连接私有 MCP 服务器。
推荐理由:官方说明了两项新能力的执行边界与网络接入方式,读者可据此判断企业私有环境下运行智能体的可行路径。
@kimmonismus@kimmonismus精选AI 评分8282 

推荐理由:原文对比了简单智能体循环与完整系统的表现,可用于判断复杂架构在什么问题上才真正必要。
@kimmonismus@kimmonismus精选AI 评分6565 代号 TrapDoor 的供应链攻击同时针对 npm、PyPI 和 Crates.io,投递 34 个恶意包,目标是加密货币、AI 和安全开发者,用于窃取钱包、SSH 密钥和云凭证。
引用Socket (@SocketSecurity)@SocketSecurityMore analysis, package details, IOCs, and GitHub-related activity here, including attacker-hosted payload/config infrastructure and PRs attempting to add .cursorrules / CLAUDE.md files to popular AI and developer projects: socket.dev/blog/trapdoor-cry…
推荐理由:该攻击把 CLAUDE.md 与 .cursorrules 配置文件当作新入口,让安全从业者了解 AI 编程助手被利用的攻击路径。
@aiDotEngineer@aidotengineer精选AI 评分6767 
推荐理由:Anthropic 工作坊聚焦让智能体持续运行数小时的工程做法,为长任务场景提供一种可对照的实现参考。
AI as Normal Technology精选AI 评分6565 AI as Normal Technology 剖析 Google 智能体 916 美元造操作系统宣称的漏洞
Sayash Kapoor 等人分析 Google 在开发者大会上发布 Gemini 3.5 Flash 和 Antigravity 2.0 时宣称的智能体团队以单一提示词、约 916.92 美元 API 费用和 2.6B tokens 造出操作系统的实验。
推荐理由:文章逐条拆解 Google 智能体造操作系统的宣称,指出单一提示词等说法缺乏关键细节,并探讨开放世界评测需要的方法规范。
Mistral AI精选AI 评分6969 Mistral 发布 Mistral Medium 3.5 并在 Vibe 和 Le Chat 推出云端远程智能体
Mistral 发布 Mistral Medium 3.5,一个 128B 稠密开源权重模型(modified MIT 许可),256k 上下文窗口,SWE-Bench Verified 得分 77.6%,τ³-Telecom 得分 91.4,自托管最少只需四块 GPU,API 定价每百万输入 token $1.5、输出 token $7.5。
推荐理由:官方同步发布模型与云端异步智能体,给出基准分数、定价和开源权重,可对照评估其编码与长程任务能力。
@swyx@swyx精选AI 评分6969 引用OpenAI Developers (@OpenAIDevs)@OpenAIDevs🥅 /goal has graduated from an experiment—for tasks big and small, Codex gets your work done. Use goal mode in the Codex app, IDE Extension, or CLI to give Codex a specific milestone, and it will keep working until it gets there, even across hours or days. You can check in and steer, and even pause Codex along the way. Pro tip: start side chats to understand the work that has been done so far without having to interrupt the main task. developers.openai.com/codex/… Video
推荐理由:原文说明 Codex 的 goal mode 已从实验转为正式功能,读者可了解它如何支撑跨小时到数天的长任务。