跳到正文

Agent 智能体

让模型自主规划、调用工具、完成多步任务的技术方向——从 Claude Code、Manus 到各家 Agent 框架与评测基准的全部动态。

当前仅显示精选新闻

最新精选

第 641–660 条 · 共 820 条
6月3日周三
  1. @alexxubyte66

    OpenAI 的数据智能体只用一个模型和 13 个工具,在 1.5 exabytes 数据和 90,000 张表上运行,设计上刻意保持简单。数据平台工程负责人 Emma Tang 介绍了其架构,以及让单一 LLM 在 90,000 张表上保持可靠的六层上下文。文章还包含 OpenAI 内部使用 Codex 的三个场景,以及五条供团队构建领域智能体参考的实践经验。

    推荐理由:多数团队会堆叠路由器、微调与多模型,OpenAI 的数据智能体却只用单模型加 13 个工具跑通 90,000 张表,形成可对比的工程取舍。

  2. @swyx74

    swyx 转发 OpenAI 公告,OpenAI 宣布扩展 Codex 插件,使其不再局限于单个工具集成。这些插件让 Codex 通过一次安装即成为特定角色的专家,无需编码,可访问 62 个常用应用和 110 项技能,覆盖销售、数据分析、创意制作、产品设计和公共股票投资等场景。

    引用OpenAI (@OpenAI)@OpenAI

    We’re making Codex more useful for your work by expanding plugins beyond individual tools. These plugins turn Codex into a specialist for a specific role with a single install, no coding required. Codex can access 62 popular apps and 110 skills for work across sales, data analytics, creative production, product design, and public equity investing. openai.com/index/codex-for-e…

    推荐理由:Codex 借插件把多款工具打包成角色化安装包,读者可据此了解其能力从编码向销售、数据分析等场景延伸。

  3. @berryxia65

    Google DeepMind 发布基于 Gemini 的多 Agent 系统 Co-Scientist,可生成、辩论并演化复杂科学问题的新假设。该系统覆盖从 idea 到验证的循环,能生成上千个假设、举办 idea 锦标赛,并让多个 Agent 互相批判精炼,再用文献、数据和搜索工具验证主张。

    引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMind

    We believe AI can be a dedicated research partner to help discover the next breakthrough. Enter Co-Scientist: our latest Gemini-based multi-agent system that can generate, debate and evolve novel hypotheses for complex scientific problems 🧵 Video

    推荐理由:原文说明了 Co-Scientist 的多 Agent 假设生成与验证流程,并指出该能力已向个人研究者开放。

  4. @eliebakouch73

    微软 MAI 技术报告因透明度受到讨论,报告显示该模型未使用合成数据或来自此前模型的蒸馏,推理、智能体行为与工具调用均在 post-training 阶段完整习得。报告给出模型各迭代阶段的精确 MFU 及对应变化,并公开完整 scaling ladder 配方,作者称这是他在同规模技术报告中第一次见到如此详细的披露。

    引用Mustafa Suleyman (@mustafasuleyman)@mustafasuleyman

    Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities. - It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks. - And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end. Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing. Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI. - Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost. All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat. Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost. Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare. Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: microsoft.ai/news/building-a…

    推荐理由:作者逐点点评微软 MAI 技术报告,读者可了解其无合成数据与蒸馏的训练取舍及 scaling ladder 的公开细节。

  5. Claude Blog61

    Anthropic 如何用 Claude 实现自助式数据分析

    Anthropic 在官方博客中介绍,其内部 95% 的业务分析查询已交由 Claude 自动完成,整体准确率约 95%。文章给出智能体分析栈的四层结构,即数据基础、真相来源、技能和验证,并归纳出概念与实体歧义、数据陈旧、检索失败三类主要错误来源。Anthropic 称未加技能时 Claude 回答分析问题的准确率不足 21%,加上技能后稳定超过 95%,部分领域可达约 99%。

    推荐理由:原文详解 Anthropic 把业务分析交给 Claude 的方法,归纳三类失败模式与分层治理思路,可迁移到其他数据团队。

  6. Claude Blog79

    Anthropic 分享 Claude Code 内部使用 skills 的经验

    Anthropic 的 Claude Code 团队分享了内部数百个 skills 的使用经验,并把它们归纳为九大类别,包括库与 API 参考、产品验证、数据获取分析、业务自动化、代码脚手架、代码质量审查、CI/CD 部署、运行手册和基础设施运维。

    推荐理由:Anthropic 把内部数百个 skills 归纳为九类并给出编写要点,可供团队搭建自己的 skills 库参考。

  7. Claude Blog63

    Anthropic Claude Code 团队如何重构 AI 原生工程组织的流程

    Anthropic Claude Code 团队负责人复盘了智能体编码成为默认工作方式后对工程流程的重写:规划从六个月路线图改成 just-in-time,遇到问题先问 Claude 并追问能否自动化,代码评审由 Claude 处理风格、缺陷和测试,人只在法律风险、信任边界与安全代码、产品判断等需要领域专长处介入。

    推荐理由:Claude Code 团队把规划、上下文获取、评审与分工按智能体编码重写,并给出可对照的前后变化与跟踪指标。

  8. @OpenAI76

    OpenAI 为 Codex 扩展 plugins,使其不再局限于单个工具。用户一次安装即可让 Codex 成为特定角色的专家,无需编码。Codex 现可访问 62 个热门应用和 110 个技能,覆盖销售、数据分析、创意生产、产品设计和公开股票投资等工作场景。

    推荐理由:原文给出可一次安装的角色化插件、接入应用与技能数量,读者可据此判断 Codex 能覆盖哪些岗位工作。

  9. @berryxia70

    OpenAI 发布 Codex Python SDK,通过 pip install openai-codex 安装后,开发者可在 Python 代码中启动线程、运行 turn、实时流式输出进度、恢复会话、传图片并精细控制 sandbox 访问权限。

    引用Vaibhav (VB) Srivastav (@reach_vb)@reach_vb

    We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install openai-codex Go build with it!!

    推荐理由:SDK 把 Codex 从浏览器里的 IDE 变成可嵌进脚本的可编程基础设施,并复用现有认证。

  10. @vista870

    OpenAI 发布 Codex Python SDK,可用 pip install openai-codex 安装,把 Codex 直接嵌入 Python 应用和工作流。该 SDK 支持启动线程、运行轮次、流式输出进度、恢复会话、传入图片和控制沙箱权限,并可复用现有 Codex 登录态。转发的作者认为这相当于把顶级编程和生图 Agent 内置到自己的代码中,并称复用 Codex 登录态最为关键。

    引用Vaibhav (VB) Srivastav (@reach_vb)@reach_vb

    We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install openai-codex Go build with it!!

    推荐理由:原文给出 Codex Python SDK 的安装命令与复用登录态方式,读者可据此把它嵌入 Python 应用和工作流。

6月2日周二
  1. Microsoft AI News60

    Microsoft AI 阐述医疗领域布局:与 Mayo Clinic 合作开发医疗前沿模型并开放 Copilot Health 预览

    Microsoft AI 发文阐述其在医疗领域的产品与研究方向,称旗下消费产品每天处理数千万条健康问题,对 50 万条 Copilot 对话的分析显示用户依赖 AI 做症状解读、检查结果解释和就医导航。

    推荐理由:微软官方阐述其在医疗领域的价值落地,读者可以据此了解 MAI DxO、Mayo Clinic 合作与 Copilot Health 的进展全貌。

  2. @MiniMax_AI73

    MiniMax 发布开源权重模型 M3,称其同时具备编码与智能体、1M 长上下文和原生多模态三项前沿能力。官方给出的基准成绩包括 SWE-Bench Pro 59.0%、Terminal Bench 2.1 66.0%、SWE-fficiency 34.8%、KernelBench Hard 28.8% 和 MCP Atlas 74.2%,并通过 MiniMax Sparse Attention 将上下文扩展至 1M。API 已在 platform.minimax.io 开放,MiniMax Code 入口为 code.minimax.io,权重与技术报告约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:官方给出 M3 在编码与智能体基准上的具体分数和 1M 上下文,读者可据此对照同类开源模型的能力边界。

  3. @MiniMax_AI68

    MiniMax 发布开源权重模型 M3,称其为首个同时具备编码、Agent 与多模态三项前沿能力的模型。M3 在 SWE-Bench Pro 得 59.0%、Terminal Bench 2.1 得 66.0%、MCP Atlas 得 74.2%,并通过 MiniMax Sparse Attention 将上下文扩展到 1M,从 Step Zero 起原生多模态。API 与 Token Plan 已在 platform.minimax.io 开放,MiniMax Code 在 code.minimax.io,权重与技术报告约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:原文列出 M3 的编码与 Agent 基准成绩及 1M 上下文能力,可对照开源模型的当前水位。

  4. Claude Blog76

    Claude Code 发布动态工作流,可为每个任务生成专属 harness

    Anthropic 在 Claude Code 中发布动态工作流,Claude 可按当前任务即时编写专属 harness,并与 Claude Opus 4.8 配合使用。工作流通过 agent() 生成子智能体,用 parallel() 和 pipeline() 组合编排,可保存为可复用工作流或通过 skill 分享。文中提到动态工作流消耗更多 token,更适合复杂高价值任务。

    推荐理由:文章给出动态工作流的六种编排模式和适用边界,读者可据此判断何时该用多智能体并行而非单上下文。

  5. @alibaba_cloud74

    阿里云发布 Qwen3.7-Plus,一个把视觉与语言统一到同一个智能体基础模型中的多模态智能体模型。官方称其支持 GUI 与 CLI 统一操作、全模态输入的编码与生产力助理,以及感知、推理、grounding 和检索增强问答的视觉智能体能力,并可跨不同智能体框架泛化。该模型现通过阿里云 Model Studio API 提供,同时给出 Qwen Studio 与博客入口。

    推荐理由:官方列出多模态智能体的能力清单与 API 入口,读者可据此判断它在自身工作流中的可用性。

  6. @kimmonismus65

    千问(Qwen)发布多模态智能体模型 Qwen3.7-Plus,将视觉与语言统一到同一智能体底座。官方列出四项能力:GUI 与 CLI 统一操作、全模态输入的编码智能体与生产力助手、具备感知推理定位与搜索增强问答的视觉智能体,以及跨智能体框架泛化,现已通过阿里云 Model Studio 开放 API(Blog:qwen.ai/blog?id=qwen3.7-plus)。转发者 @kimmonismus 认可其多模态表现,但质疑官方为何将自家模型与 GPT-5.4、Opus 4.6 对比。

    引用Qwen (@Alibaba_Qwen)@Alibaba_Qwen

    👏👏 Introducing Qwen3.7-Plus — a multimodal agent model that unifies vision and language into one versatile agent foundation. ✅ Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks ✅ Versatile coding agent & productivity assistant with full-modality input ✅ Visual Agent: perception, reasoning, grounding, and search-augmented QA ✅ Cross-harness generalization across diverse agent frameworks One model. Sees, thinks, codes, acts.🙌🙌 Now available via API on Alibaba Cloud Model Studio. Try it — let us know what you build.😎 🔗🔗⬇️⬇️ Blog:qwen.ai/blog?id=qwen3.7-plus Qwen Studio:chat.qwen.ai/?models=qwen3.7… API:modelstudio.console.alibabac…

    推荐理由:官方把 Qwen3.7-Plus 与 GPT-5.4、Opus 4.6 并列对比,读者可据此观察多模态智能体模型的竞争位置。

  7. @kimmonismus67

    NVIDIA 发布 DGX Station for Windows 桌面级 AI 超算,搭载 GB300 Grace Blackwell Ultra 桌面超级芯片,最高 748GB 一致内存、20 petaflops FP4 算力,可本地运行最高 1 万亿参数模型,Q4 出货。

    引用NVIDIA Newsroom (@nvidianewsroom)@nvidianewsroom

    Introducing NVIDIA DGX Station for Windows, the world's most powerful deskside AI supercomputer with Windows powered by NVIDIA GB300. ✅ Run frontier AI models with up to 1 trillion parameters locally ✅ Build and run secure AI agents on Windows with NVIDIA OpenShell ✅ Built by @ASUS, @Dell, @GIGABYTE, @HP, @msigaming, and @Supermicro #NVIDIAGTC nvda.ws/3RIkjpc

    推荐理由:原文列出 GB300 桌面超级芯片的完整规格与出货时间,可据此判断本地运行前沿模型与智能体的硬件门槛。

  8. @kilocode67

    MiniMax 发布 M3 模型,官方称其为首个同时整合编程、Agent 与多模态三项前沿能力的开放权重模型。官方报告其在 SWE-Bench Pro 取得 59.0%、Terminal Bench 2.1 取得 66.0%、SWE-fficiency 取得 34.8%、KernelBench Hard 取得 28.8%、MCP Atlas 取得 74.2%,并通过 MiniMax Sparse Attention 将上下文扩展到 1M。模型从第一步起原生多模态,权重与技术报告预计约 10 天后发布,API 与 MiniMax Code 入口已开放。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:官方给出 M3 的编程与 Agent 基准数据及 1M 上下文能力,可据此观察开源权重模型的前沿能力边界。

6月1日周一
  1. @kimmonismus74

    MiniMax 发布 M3 开放权重模型,在 SWE-Bench Pro 上得分 59%,略高于 GPT-5.5 的 58.6%,高于 Gemini 3.1 Pro 的 54.2%,编码上仍落后 Opus 4.7。该模型支持 1M token 上下文并原生多模态,在 BrowseComp 上以 83.5% 领先 Opus 4.7,每 token 成本约为 GPT-5.5 的 1/12。权重和完整技术报告预计约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:原文给出了 M3 与 GPT-5.5、Opus 4.7 的基准对比和成本差距,读者可据此判断开放权重模型的能力位次。