Anthropic 发布 Claude Opus 4.8,并上线动态工作流与 effort 控制
Anthropic 发布 Claude Opus 4.8,在编码、智能体与推理等基准上较 Opus 4.7 有提升,常规使用价格与 Opus 4.7 相同,为每百万输入 token 5 美元、输出 25 美元。
推荐理由:官方给出与前代及 GPT-5.5、Gemini 3.1 Pro 的基准对比,读者可据此判断它在编码与智能体任务中的能力位置。
AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。
当前仅显示精选新闻Anthropic 发布 Claude Opus 4.8,在编码、智能体与推理等基准上较 Opus 4.7 有提升,常规使用价格与 Opus 4.7 相同,为每百万输入 token 5 美元、输出 25 美元。
推荐理由:官方给出与前代及 GPT-5.5、Gemini 3.1 Pro 的基准对比,读者可据此判断它在编码与智能体任务中的能力位置。
Anthropic 为 Claude Code 推出动态工作流,Claude 可动态编写编排脚本,在单次会话中并行运行数十到数百个子智能体,并在结果交付前自行校验。
推荐理由:官方给出了并行子智能体编排的能力边界与 Bun 重写中的实际用例,便于判断它能否承接超长任务。
1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.
推荐理由:作者从模型多样性、GTM 与企业级验证等角度拆解 Cognition 的护城河,可结合其融资数据理解这轮投资逻辑。
Anthropic 发布 Claude Opus 4.8,为其最强模型,在 Claude API、Amazon Bedrock、Google Cloud 和 Microsoft Foundry 上默认支持 1M token 上下文窗口,最大输出 128k token。
推荐理由:原文完整列出 1M 上下文、采样参数限制和各端可用性等迁移要点,接入方可以直接对照判断升级路径。
Anthropic 发布指南,介绍如何用 Claude Opus 通过威胁建模、沙箱、发现、验证、分诊、修补六个步骤查找并修复源码漏洞。文中称截至 2026 年 5 月 22 日其开源软件扫描已披露 1596 个漏洞,其中 97 个已修补,瓶颈已从发现转向验证、分诊与修补。
推荐理由:文章把漏洞发现到修补拆成六步方法并附可复用 skill 与仓库,瓶颈已转向验证与修补。
🥷 Engineering habits you already know, turned into skills AI agents can run. Waza absorbed a mass of real project lessons recently. Now just as sharp for Mac native apps, CLI tools, and Rust as it is for web. Supports Claude Code, Codex, Cursor, and Pi as agent runtimes. Reviews your CLI like a shipped product. Debugs "works in source tree, breaks after install." Sweeps sibling instances after every fix. Blocks "fixed" until runtime evidence is verified. 25 anti-patterns, destructive command safety, treats fetched content as untrusted data. 8 skills, no framework, no telemetry. Your superpower prompt collection can be uninstalled. Too heavy. github.com/tw93/Waza
推荐理由:原文给出把 Skill 写入 Agent.md 的提示词写法,读者可据此改善技能触发与应用效果。
代号 TrapDoor 的供应链攻击同时针对 npm、PyPI 和 Crates.io,投递 34 个恶意包,目标是加密货币、AI 和安全开发者,用于窃取钱包、SSH 密钥和云凭证。
More analysis, package details, IOCs, and GitHub-related activity here, including attacker-hosted payload/config infrastructure and PRs attempting to add .cursorrules / CLAUDE.md files to popular AI and developer projects: socket.dev/blog/trapdoor-cry…
推荐理由:该攻击把 CLAUDE.md 与 .cursorrules 配置文件当作新入口,让安全从业者了解 AI 编程助手被利用的攻击路径。
Mistral 发布 Mistral Medium 3.5,一个 128B 稠密开源权重模型(modified MIT 许可),256k 上下文窗口,SWE-Bench Verified 得分 77.6%,τ³-Telecom 得分 91.4,自托管最少只需四块 GPU,API 定价每百万输入 token $1.5、输出 token $7.5。
推荐理由:官方同步发布模型与云端异步智能体,给出基准分数、定价和开源权重,可对照评估其编码与长程任务能力。
🥅 /goal has graduated from an experiment—for tasks big and small, Codex gets your work done. Use goal mode in the Codex app, IDE Extension, or CLI to give Codex a specific milestone, and it will keep working until it gets there, even across hours or days. You can check in and steer, and even pause Codex along the way. Pro tip: start side chats to understand the work that has been done so far without having to interrupt the main task. developers.openai.com/codex/… Video
推荐理由:原文说明 Codex 的 goal mode 已从实验转为正式功能,读者可了解它如何支撑跨小时到数天的长任务。
推荐理由:官方给出 Qwen3.7-Max 在编码智能体方向的能力定位和一个月五折的上线优惠,读者可据此判断是否值得试用。
It’s Codex Thursday, and yes, we have updates for you. First up: Appshots, a new way to bring the context of what you’re working on into Codex. On your Mac, press Command-Command to attach your app window to a Codex thread. Codex gets both a screenshot and text from the window, including content beyond what’s visible onscreen. Appshots are available across plans on Mac, with enterprise access coming soon. Video
推荐理由:介绍了快捷截图、/goal 与浏览器注释等更新的实际用法,可据此了解 Codex 工作流的变化。
推荐理由:Codex 新增的 Goal mode 让智能体可围绕一个目标持续工作数小时甚至数天,读者可据此判断长任务自动化的可用边界。
推荐理由:官方说明 Codex 可在锁屏息屏时从手机操控 Mac 应用,读者可据此判断远程编程工作流的可用性。
🥅 /goal has graduated from an experiment—for tasks big and small, Codex gets your work done. Use goal mode in the Codex app, IDE Extension, or CLI to give Codex a specific milestone, and it will keep working until it gets there, even across hours or days. You can check in and steer, and even pause Codex along the way. Pro tip: start side chats to understand the work that has been done so far without having to interrupt the main task. developers.openai.com/codex/… Video
推荐理由:Codex 的 goal 模式从实验转为正式功能,任务可跨小时或跨天持续推进,并附有启用命令。
推荐理由:goal mode 由实验功能转为正式可用,读者可据此了解 Codex 在跨小时任务上的执行与介入方式。
The new Qwen3.7-Max from @Alibaba_Qwen is live on OpenRouter. The flagship of the Qwen3.7 series, built for agent-centric work: coding, office and productivity tasks, and long-horizon autonomous execution. Big jumps in coding and agent benchmarks over Qwen3.6, with explicit prompt caching for repeated context.
推荐理由:Qwen3.7-Max 已在 OpenRouter 上线,面向编码与办公的智能体场景,读者可了解这一旗舰版本的定位。
In the next version of Claude Code: run /usage to see a breakdown of which Skills, Agents, MCPs, and Plugins are using your tokens CLI today, coming to Desktop next
推荐理由:作者贴出自己的用量数据,说明 /usage 能把 token 消耗归因到具体子代理和 MCP,读者可据此排查隐性开销。
📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get things done: 🧑💻 Coding agent, end to end. Frontend prototypes, multi-file refactors, real debugging — nails it. 🗂️ A reliable office and productivity assistant. Get your work done through MCP integrations and multi-agent orchestration. ⏱️ Long-horizon autonomy. 35 hours straight on a kernel optimization task — 1,000+ tool calls, zero hand-holding. 🔌 Scaffold-agnostic. Claude Code, OpenClaw, Qwen Code, or your own stack. Consistent reliability everywhere. API's up on Alibaba Model Studio. You can also take it for a spin on Qwen Studio. Go build something wild!🏃🏃♂️ 📖 Blog: qwen.ai/blog?id=qwen3.7 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.7… ⚡️ API:modelstudio.console.alibabac…
推荐理由:原文列出端到端编码、MCP 集成与长时间自主运行等能力,可用以判断该旗舰模型在 Agent 场景中的定位。
推荐理由:Qwen3.7-Max 作为千问旗舰上线 OpenRouter,面向智能体任务,并给出相对 Qwen3.6 的基准变化。
Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past releases @cursor_ai has released Composer 2.5, the latest model in its Composer line. Composer 2.5 scored 62 on our Coding Agent Index, a 14 point gain over Composer 2 (48). This puts it in third place of our tested agents, behind only Claude Opus 4.7 (max) in Claude Code (66) and GPT-5.5 (xhigh reasoning) in Codex (65). These cost $4.10 and $4.82 per task respectively, ~10x the cost of Composer 2.5 Fast ($0.44) and ~60x the cost of Composer 2.5 standard ($0.07). Key results for Composer 2.5 in Cursor CLI: ➤ Cost-quality Pareto frontier: At $0.07 (standard) and $0.44 (Fast) per task, Composer 2.5 is cheaper than every other agent scoring above 60 on the Index. Medium-effort peers cost $1.24–$2.21 per task; higher-effort variants land 3-4 points above at $4.10–$4.82 ➤ Per-benchmark gains vs Composer 2: +35 points on SWE-Bench-Pro-Hard-AA (12% → 47%), +2 points on Terminal-Bench v2 (64% → 66%), and +3 points on SWE-Atlas-QnA (69% → 72%). At 47%, Composer 2.5's score on SWE-Bench-Pro-Hard-AA is comparable to Claude Opus 4.7 (max) in Claude Code ➤ Among the fastest coding agents: Composer 2.5 Fast runs at an average wall time of 6.7 minutes per task, the third-fastest agent on the Artificial Analysis Coding Agent Index, behind only Claude Opus 4.7 (medium) in Claude Code (5.8m) and GPT-5.5 (medium) in Cursor CLI (6.2m) ➤ Fast mode enables better responsiveness at 6x pricing: Fast runs 30% faster than standard Composer 2.5, but is ~6x the cost per task ($0.44 vs $0.07). Token pricing is 6x higher for Fast: $3.00/$15.00 vs $0.50/$2.50 per million input/output tokens Model details: ➤ Base model: Continued training on @Kimi_Moonshot's open weights Kimi K2.5 as with Composer 2, with Cursor reporting ~85% of total compute from its own additional training and reinforcement learning ➤ Pricing: $0.50/$2.50 per million input/output tokens for the standard variant; $3.00/$15.00 for the Fast variant (the default in Cursor) ➤ Available exclusively in Cursor: both Cursor IDE and Cursor CLI, an externally accessible API is not available Congratulations @cursor_ai and @mntruell on the impressive release!
推荐理由:推文用每任务成本对比 Composer 2.5 与两个更高分编码智能体,读者可据此权衡编码任务上的性能与花费。