推荐理由:原文列出了 Grok Bot 面向 Cursor Pro+、Teams 与 SuperGrok Plus 的开放范围,读者可据此判断自己能否直接使用。
Agent 智能体
全部主题让模型自主规划、调用工具、完成多步任务的技术方向——从 Claude Code、Manus 到各家 Agent 框架与评测基准的全部动态。
当前仅显示精选新闻最新精选
第 381–400 条 · 共 874 条@AYi_AInotes@AYi_AInotes精选AI 评分6868 
@AYi_AInotes@AYi_AInotes精选AI 评分7272 墨问西东创始人池建强体验 DeepSeek Harness 一天后,叫停刚完成两个多月重构的客户端开发,开始论证全面迁移。
推荐理由:借一位创始人停掉两年自研客户端的决定,呈现插件化 Agent 底座对应用开发与迁移选择的影响。
Claude Blog精选AI 评分6969 Anthropic 发布 AI-Native SDLC 手册,梳理 Claude 在六个研发阶段的落地方式
Anthropic 发布 AI-Native SDLC playbook,把软件开发流程拆成 Plan、Design、Build、Test、Deploy、Maintain 六个阶段,逐阶段给出具体做法、治理考量与度量指标。
推荐理由:Anthropic 公开其内部把 Claude 嵌入研发六个阶段的做法,含各阶段产物与治理门槛,便于团队对照改造自身流程。
@omarsar0@omarsar0精选AI 评分6767 引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:官方基准对比显示该实验模型在多模态智能体任务上接近 Opus-4.8,读者可据此了解多模态智能体的当前水平。
@AYi_AInotes@AYi_AInotes精选AI 评分7272 

引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:逐条列出这个实验性多模态模型的单图 token 上限、分级定价与免费 Files API 额度,可供评估视觉 Agent 的成本与接入方式。
@kimmonismus@kimmonismus精选AI 评分6767
引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:基准对比显示 Flash 小模型在多模态智能体评测上已接近 Opus-4.8,可供判断轻量模型的能力边界。
@testingcatalog@testingcatalog精选AI 评分7171
引用@deepseek_ai@deepseek_aiMultimodal API support 🔌 🔹 Set model='deepseek-v4-flash-vision-exp' 🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing 🔹 Supports Chat Completions, Messages & Responses 🔹 Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API. Docs: https://t.co/USZ5gZ3wWB 3/n
推荐理由:材料给出新模型与 Opus 4.8 及自家文本版本的基准对比和计费细节,便于判断实际可用性。
IT Home精选AI 评分8080 OpenAI 开源 Codex Harness,放出 CLI、SDK 与 app-server 三类组件
OpenAI 以 Apache-2.0 许可开源驱动 Codex 的底层执行框架 Harness,并放出三类组件:运行自动化流水线的 codex exec CLI、支持 TypeScript 和 Python 的 Codex SDK,以及通过 JSON-RPC 连接本地 Codex 进程的 app-server。
推荐理由:开源的三类组件与 Harness 调优带来的基准变化,可供开发者评估把智能体嵌入自有产品的落地方式。
Together AI Blog精选AI 评分7171 Together AI 实测 GLM-5.3 对比 GPT-5.6 Sol:DeepSWE 上的成本、编码与路由策略
Together AI 在 DeepSWE 全部113个任务上以每配置4次试验、共904次rollout实测 GLM-5.3 (max) 与 GPT-5.6 Sol (max)。
推荐理由:原文基于904次rollout的实测数据对比两款模型的成本、pass@k和失败模式,还给出可复用的级联路由方案。
机器之心 · 微信公众号精选AI 评分7777 Anthropic 拟最早8月底提交 S-1,IPO 目标估值1.5-2万亿美元
据彭博社报道,Anthropic 正加速公开上市进程,计划最早于今年8月底向 SEC 公开提交 S-1 招股说明书,目标估值1.5-2万亿美元,募资规模有望持平或超越 SpaceX 创下的750亿美元纪录,预计最快今年第四季度挂牌。
推荐理由:材料给出 Anthropic 的上市时间表、募资体量和营收运行率数据,读者可据此看清前沿模型实验室登陆公开市场的节奏与规模。
@ClaudeDevs@ClaudeDevs精选AI 评分7070 
推荐理由:这些能力从测试转为正式可用,开发者可在没有 API 的应用里做自动化,并用版本化技能和可复用文件搭建托管智能体。
Claude Blog精选AI 评分7272 Claude 平台 computer use、Skills API 和 Files API 正式可用
Anthropic 宣布 computer use、Skills API 和 Files API 在 Claude Platform 正式可用,computer use 同时新增 browser use 工具,除截图外还读取页面结构,让智能体定位具体字段或按钮。
推荐理由:正式版给出了多动作执行、浏览器工具和文件存储的具体变化,可据此判断现有智能体工作流的改造空间。
xAI News精选AI 评分7070 xAI 将 Grok Build 开放至网页、iOS 和 Android
xAI 将 Grok Build 从 7 月起仅限 SuperGrok Heavy 的 Early Beta 开放给所有用户,网页、iOS 和 Android 均可使用,只需在对话中描述应用、游戏、网站或仪表盘即可生成可运行版本。
推荐理由:Grok Build 从仅限 SuperGrok Heavy 的早期测试转向全量开放,并补齐发布分享、自有域名与 GitHub 导出等交付环节。
DeepSeek API updates精选AI 评分7070 DeepSeek 发布实验版多模态模型 DeepSeek-V4-Flash-Vision-Exp
DeepSeek 在 API 平台上线实验性多模态视觉理解模型 DeepSeek-V4-Flash-Vision-Exp,通过设置 model='deepseek-v4-flash-vision-exp' 即可调用。
推荐理由:官方公布了多模态 Agent 基准成绩和调用方式,读者可据此判断其视觉 Agent 能力在现有模型中的位置。
@rohanpaul_ai@rohanpaul_ai精选AI 评分6565 
推荐理由:论文给出技能库持久化风险的量化证据,并附检测基准与修复方案,可迁移到智能体安全评估。
Mistral AI精选AI 评分6666 Mistral 发布 Agentic Search,多步检索在 FinanceBench 上将准确率从 26.7% 提升至 86%
Mistral 发布 Agentic Search,通过 search、open、navigate、read、grep 五个工具让模型多步检索、翻阅并验证复杂文档,可通过 Mistral Search Toolkit 及 Studio 和 Vibe 中的 Libraries 使用,无需微调。
推荐理由:官方给出多步检索循环的完整工具设计和两个基准的具体数字,读者可据此评估它替代一次性 RAG 的适用场景。
@kimmonismus@kimmonismus精选AI 评分7171 
推荐理由:原文给出 Asana 大型前端迁移的工时与成本对比,可作为判断智能体承接遗留代码改造的参考。
@testingcatalog@testingcatalog精选AI 评分7272 
推荐理由:Cursor 让子智能体在独立虚拟机中运行并支持定时任务,读者可据此了解云端编码智能体的隔离与自动化方向。
@gdb@gdb精选AI 评分6565 引用@OpenAIDevs@OpenAIDevsTeams are using the open-source Codex harness to bring agents into the tools they already use, from internal apps to operations dashboards. Their applications control the interface, context, tools, and approvals while the harness handles the agent loop. https://t.co/shi2vwqxh4
推荐理由:税务试点把 Codex 用在编码之外,配合开源 harness 展示了智能体接入自有产品的一条路径。
@rohanpaul_ai@rohanpaul_ai精选AI 评分6767 
推荐理由:泄露的投资人信给出 OpenRouter 的 token 消耗增速,可看到支付公司如何把智能体当作经济参与者来定价。