跳到正文

Agent 智能体

让模型自主规划、调用工具、完成多步任务的技术方向——从 Claude Code、Manus 到各家 Agent 框架与评测基准的全部动态。

当前仅显示精选新闻

最新精选

第 701–720 条 · 共 798 条
5月21日周四
  1. @kimmonismus69

    Cursor 发布 Composer 2.5,在 Artificial Analysis 编码智能体指数上得分 62,比上一代 Composer 2 提升 14 分,位列第三,仅次于 Claude Opus 4.7(max)的 66 分和 GPT-5.5(xhigh)的 65 分。标准版每任务成本 0.07 美元、Fast 版 0.44 美元,而上述两款更高分模型分别约为 4.10 和 4.82 美元。该模型仅在 Cursor IDE 和 Cursor CLI 提供,无外部 API,基于 Kimi K2.5 继续训练;推文作者认为性能略好却贵 60 倍已不再划算。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past releases @cursor_ai has released Composer 2.5, the latest model in its Composer line. Composer 2.5 scored 62 on our Coding Agent Index, a 14 point gain over Composer 2 (48). This puts it in third place of our tested agents, behind only Claude Opus 4.7 (max) in Claude Code (66) and GPT-5.5 (xhigh reasoning) in Codex (65). These cost $4.10 and $4.82 per task respectively, ~10x the cost of Composer 2.5 Fast ($0.44) and ~60x the cost of Composer 2.5 standard ($0.07). Key results for Composer 2.5 in Cursor CLI: ➤ Cost-quality Pareto frontier: At $0.07 (standard) and $0.44 (Fast) per task, Composer 2.5 is cheaper than every other agent scoring above 60 on the Index. Medium-effort peers cost $1.24–$2.21 per task; higher-effort variants land 3-4 points above at $4.10–$4.82 ➤ Per-benchmark gains vs Composer 2: +35 points on SWE-Bench-Pro-Hard-AA (12% → 47%), +2 points on Terminal-Bench v2 (64% → 66%), and +3 points on SWE-Atlas-QnA (69% → 72%). At 47%, Composer 2.5's score on SWE-Bench-Pro-Hard-AA is comparable to Claude Opus 4.7 (max) in Claude Code ➤ Among the fastest coding agents: Composer 2.5 Fast runs at an average wall time of 6.7 minutes per task, the third-fastest agent on the Artificial Analysis Coding Agent Index, behind only Claude Opus 4.7 (medium) in Claude Code (5.8m) and GPT-5.5 (medium) in Cursor CLI (6.2m) ➤ Fast mode enables better responsiveness at 6x pricing: Fast runs 30% faster than standard Composer 2.5, but is ~6x the cost per task ($0.44 vs $0.07). Token pricing is 6x higher for Fast: $3.00/$15.00 vs $0.50/$2.50 per million input/output tokens Model details: ➤ Base model: Continued training on @Kimi_Moonshot's open weights Kimi K2.5 as with Composer 2, with Cursor reporting ~85% of total compute from its own additional training and reinforcement learning ➤ Pricing: $0.50/$2.50 per million input/output tokens for the standard variant; $3.00/$15.00 for the Fast variant (the default in Cursor) ➤ Available exclusively in Cursor: both Cursor IDE and Cursor CLI, an externally accessible API is not available Congratulations @cursor_ai and @mntruell on the impressive release!

    推荐理由:推文用每任务成本对比 Composer 2.5 与两个更高分编码智能体,读者可据此权衡编码任务上的性能与花费。

  2. @kimmonismus72

    阿里发布旗舰模型 Qwen3.7-Max,面向智能体场景,官方称其在一次内核优化任务中自主运行 35 小时、发起 1,158 次工具调用,并在单个注意力内核上取得 10 倍加速,模型已上线阿里云 Model Studio,也可在 Qwen Studio 试用。

    引用Qwen (@Alibaba_Qwen)@Alibaba_Qwen

    📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get things done: 🧑‍💻 Coding agent, end to end. Frontend prototypes, multi-file refactors, real debugging — nails it. 🗂️ A reliable office and productivity assistant. Get your work done through MCP integrations and multi-agent orchestration. ⏱️ Long-horizon autonomy. 35 hours straight on a kernel optimization task — 1,000+ tool calls, zero hand-holding. 🔌 Scaffold-agnostic. Claude Code, OpenClaw, Qwen Code, or your own stack. Consistent reliability everywhere. API's up on Alibaba Model Studio. You can also take it for a spin on Qwen Studio. Go build something wild!🏃🏃‍♂️ 📖 Blog: qwen.ai/blog?id=qwen3.7 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.7… ⚡️ API:modelstudio.console.alibabac…

    推荐理由:作者把 35 小时自主优化的传播印象与实际范围区分开,并单独讨论智能体能力泛化这一论断。

  3. @alibaba_cloud66

    阿里云披露 Qwen3.7 的自主进化实验:在约 35 小时连续自主执行中,模型完成 432 次 kernel 评估、跨 1158 次工具调用,独立编写、编译、剖析并迭代改进 Extend Attention Kernel。在多种工作负载下,该 kernel 相对 Triton 参考实现取得 10.0x 几何平均加速,更多细节见 qwen.ai/blog?id=qwen3.7。

    推荐理由:官方披露 Qwen3.7 连续自主执行约 35 小时优化注意力 kernel 的过程,可作为观察自主编码智能体能力的参照。

  4. @alibaba_cloud70

    阿里云发布旗舰模型 Qwen3.7-Max,定位为面向智能体时代的基础模型,API 已在 Model Studio 上线。官方称其可端到端完成编码任务,包括前端原型、多文件重构与调试,并可通过 MCP 集成和多智能体编排承担办公与生产力助手工作。

    推荐理由:官方列出编码、长任务与多种脚手架兼容能力,读者可据此判断这款旗舰模型在智能体工作流中的定位。

  5. @googleaidevs67

    Google 为智能体设计工具 Stitch 推出多项更新,现可实时流式生成设计并当场编辑、接收交互反馈,支持直接导入代码库或 Design.md 以沿用既有生产组件,还能生成动态界面并将项目导出为可分享的线上 URL。这些更新已在全球上线,地址为 stitch.withgoogle.com。

    推荐理由:官方列出 Stitch 从实时预览、代码库导入到导出线上 URL 的四项更新,可据此判断原型到部署链路的变化。

5月20日周三
  1. @berryxia73

    Google 发布 Gemini 3.5 Flash,Artificial Analysis 测试显示其 Intelligence Index 为 55 分,比 Gemini 3 Flash 高 9 分,超过 Grok 4.3 和 Claude Sonnet 4.6,输出速度超 280 tokens/s,比上一代快 70%,幻觉率从 92% 降到 61%。

    引用Berryxia.AI (@berryxia)@berryxia

    兄弟们! 今天已经可以在ZenMux上免费体验Gemini 3.5 Flash 了! 我第一时间用它跑了那个经典的「AI模型递归二叉树生长测试」. 同一个 Prompt ,不同模型画出的树形态完全不一样。(见视频-Prompt见评论区) Gemini 3.5 Flash 从输入提示词到生成完整 HTML 动画网页(树干慢慢长出、分支递归展开、最后随风摇摆),全程只用了 77.56 秒! 整体效果非常惊艳:树形态自然优雅、生长动画丝滑、视频和内容呈现都顶级! 熟悉的老朋友都知道,ZenMux 每次新模型都是 ZeroDelay 首发. Google I/O 2026 今天刚发布,现在立刻就能通过 API 调用! 还有免费额度可以白嫖~ 速度是真的没话说,还完美保留了旗舰级模型的能力。 专为 Agent 设计,在 MCP Atlas、Toolathlon、Finance Agent 等多项榜单直接拿下第一! 多模态理解也极强:MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,全面超越上一代 Gemini 3.1 Pro。 完全兼容主流 API 格式,无需改动现有工具链。 支持按量计费 + Builder 套餐。 👇 直接体验 正式版 → zenmux.ai/google/gemini-3.5-… 免费试用 → zenmux.ai/google/gemini-3.5-… Video

    推荐理由:原文用基准与定价的对比说明 Flash 系列的定位变化,读者可据此重新评估轻量模型的成本预期。

  2. 微信公众号(账号未识别)85

    Anthropic 发布 Claude Managed Agents,提供云托管 Agent 编排 API

    Anthropic 发布 Claude Managed Agents,一套用于构建和部署云托管 AI Agent 的可组合 API,核心是经调优的 Harness 编排引擎。架构上将 Session、Harness、Sandbox 三层解耦,凭证不进入沙箱,官方称 p50 TTFT 下降约 60%。定价为标准 token 费率加每 Session 活跃运行时间 $0.08/小时,空闲时间不计费。

    推荐理由:文章拆解了 Managed Agents 的 Harness 解耦架构与计费方式,可据此理解 Anthropic 转向 Agent 基础设施的路径。

  3. @berryxia72

    Gemini 3.5 Flash 已在 ZenMux 上线并提供免费试用,作者实测用它从提示词生成完整 HTML 递归树生长动画,全程耗时 77.56 秒。该模型在 MCP Atlas、Toolathlon、Finance Agent 等榜单拿下第一,MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,并兼容主流 API 格式。

    推荐理由:作者用递归树动画实测 Gemini 3.5 Flash 的生成速度,并列出其在 Agent 榜单与多模态基准上的成绩。

  4. @berryxia79

    Google DeepMind 发布 Gemini 3.5 Flash,Artificial Analysis 预发布测试显示其 Intelligence Index 得 55 分,比 Gemini 3 Flash 高 9 分。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in its Flash family, which has traditionally has offered faster, lower-cost alternatives to Gemini Pro models. Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash, driven primarily by agentic performance gains and hallucination reduction. It achieves speeds of over 280 output tokens/s, but higher token usage and token pricing make it over 5x more costly to run the Intelligence Index than Gemini 3 Flash, and 75% more costly than Gemini 3.1 Pro. Gemini 3.5 Flash is $1.50/1M input and $9/1M output tokens, Gemini 3 Flash was $0.5/$3 per 1M input/output tokens, a 3x increase. The rest of the increase was driven by higher token usage when running our benchmarks Key results for Gemini 3.5 Flash with ‘high’ thinking level: ➤ 9 point Intelligence Index improvement: Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash. This places it ahead of Grok 4.3 (high, 53) and Claude Sonnet 4.6 (max, 52). The model improves across nearly all evaluations, with the largest gains coming from agentic evaluations and AA-Omniscience (knowledge and hallucination). On AA-Omniscience, Gemini 3.5 Flash improves by 11 points, driven primarily by reduced hallucinations, with its hallucination rate falling to 61%, a 31 point decrease compared to Gemini 3 Flash ➤ Agentic capability improvements: Gemini 3.5 Flash improves substantially over Gemini 3 Flash across our agentic evaluations, in both GDPval-AA (real-world agentic tasks) and Tau2-Bench Telecom (agentic tool use). Its GDPval-AA result is especially notable, achieving an Elo of 1656, well ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), and just behind GPT-5.4 (xhigh, 1674). This represents a meaningful step forward for Google in agentic performance, which has historically been a relative weakness for Gemini models ➤ Speed-intelligence frontier: Gemini 3.5 Flash achieves speeds of over 280 output tokens per second, ~70% faster than Gemini 3 Flash and models such as gpt-oss-120b and GPT-5.4 mini (xhigh). With its 55 Intelligence Index score, this places Gemini 3.5 Flash on the speed-intelligence Pareto frontier alongside Gemini 3.1 Pro and Gemini 3.1 Flash-Lite, reinforcing Google’s strength in models balancing speed and intelligence ➤ 5.5x increase in cost to run: Gemini 3.5 Flash costs $1,552 to run the Artificial Analysis Intelligence Index, 5.5x more than Gemini 3 Flash and 75% more than Gemini 3.1 Pro. This is driven by increases in both token usage and token prices. Output token usage is broadly unchanged from Gemini 3 Flash (73M vs. 72M), but input token usage increases significantly, driven primarily by an increase in the number of turns in agentic evaluations. Gemini 3.5 Flash is priced 3x higher than Gemini 3 Flash at $1.50/$9.00 per 1M input/output tokens, with a 90% discount for cached input tokens ➤ Google continues to lead multimodal performance: Gemini 3.5 Flash is multimodal, supporting image, video, and speech input alongside text. This differs from many proprietary models, including Claude Opus 4.7, Grok 4.3, and GPT-5.5, which support image input only. In our multimodal evaluation, MMMU-Pro, Gemini 3.5 Flash scores 84% - the highest score recorded. This puts models from Google in the top two spots, with Gemini 3.1 Pro scoring 82% Key model details: ➤ Context window: Retains the same 1M context window as Gemini 3 Flash ➤ Multimodality: Text, image, video and speech input with text output only ➤ Pricing: $1.50/$9.00 per million input/output tokens, with a 90% discount for cached input tokens Congratulations @GoogleDeepMind , @sundarpichai and @demishassabis on the great release!

    推荐理由:借 Artificial Analysis 的预发布基准,可以看到 Gemini 3.5 Flash 在智能与速度上的提升及其成本代价。

  5. @OpenRouter79

    Google DeepMind 发布 Gemini 3.5 模型家族,称其将前沿智能与现实世界行动结合,3.5 Flash 是该家族发布的第一款模型,官方称其为面向智能体与编码的最强模型。OpenRouter 转发该消息,并附上阅读模型详情的链接。

    引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMind

    Introducing Gemini 3.5: our newest family of models combining frontier intelligence with real-world action. The first release is 3.5 Flash, our strongest model yet for agents and coding 🧵

    推荐理由:Google DeepMind 公布 Gemini 3.5 家族,3.5 Flash 主打智能体与编码,可看到新模型的能力侧重。

  6. @GeminiApp69

    Google 发布个人 AI 智能体 Gemini Spark,可在后台自主执行任务,即使手机和笔记本电脑处于关机状态也能运行。该智能体由用户主动开启,并在执行重大操作前会先与用户确认。

    引用Google Gemini (@GeminiApp)@GeminiApp

    Gemini Spark is your new 24/7 personal AI agent. Give it a task and it works autonomously in the background, even if your phone and laptop are turned off. You choose to turn it on and it's designed to check with you before taking major actions. #GoogleIO

    推荐理由:原文交代了 Gemini Spark 的后台常驻运行方式与重大操作前确认机制,便于判断个人智能体的授权边界。