#Agent
#Agent
今日 81 条
elvis@omarsar0AI 评分4848
Hugging Face Daily PapersAI 评分4747 AutoGUIWorld:用图像生成器作为 GUI 智能体的视觉世界模型
AutoGUIWorld 是一个数据生成框架,结合图像生成器的视觉先验与规划器的任务知识,无需部署或运行软件环境即可合成 GUI 交互轨迹。它从操作系统上下文、视觉外观和界面状态的结构化规格中采样初始场景,生成 79,266 条覆盖 Ubuntu、Windows、macOS 和 Chrome 的空间标注步级训练样本。
Hugging Face Daily PapersAI 评分3737 FloWright:用工作流优化工作流,小模型性能提升最高 7.41%
针对多智能体工作流训练只优化生成器、其余智能体固定的问题,研究者提出 FloWright,通过分层、结构感知的奖励范式让一个角色自我进化、两个及以上角色协同进化,无需额外模型、标签或执行。
ginobefun@hongming731AI 评分2626LangChain 在 Open SWE 智能体 Harness 内构建模型路由器,按任务分类选择最低成本合适模型,将编码线程中位数成本降低 64%,质量变化可忽略不计。
引用ginobefun@hongming731https://x.com/i/article/2105830283580989440
elvis@omarsar0AI 评分5050引用Tapa Ghosh@semiDLExcited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post
Hacker News popular via buzzing.ccAI 评分6363 arXiv 论文提出 Context Language Models,让模型原生管理自身上下文
Rulin Shao 等人在 arXiv:2609.37725 提出 Context Language Models(CLMs),把上下文当作文件、由模型自由更新,从而原生管理自身上下文,并可自然扩展到多智能体共享文件式上下文。
Latent SpaceAI 评分5252 Latent Space 访谈 MIT 的 Alex Zhang:RLM、harness 设计与研究品味
Latent Space 播客访谈 MIT 博士生 Alex Zhang,围绕其 Recursive Language Models(RLM)研究展开。
Artificial Analysis@ArtificialAnlys精选AI 评分7070
推荐理由:原文给出三大新模型组合在 Coding Agent Index 的得分与单任务成本对比,读者可以据此在性能和价格之间做选型权衡。
OpenRouter Announcements精选AI 评分6464 OpenRouter 详解六大 Agent 框架的工具调用 Schema 处理,并提出在 API 层统一格式
OpenRouter 比较了 LangChain、LangGraph、CrewAI、OpenAI Agents SDK、Claude Agent SDK、Microsoft Agent Framework 和 Google ADK 如何定义工具 Schema 并在不同提供商的 wire format 之间做翻译。
推荐理由:原文逐一拆解六大框架的工具调用格式翻译位置,并给出在 API 层统一格式的可行做法,便于开发者选型前对齐自己的技术栈。
OpenRouter AnnouncementsAI 评分5959 OpenRouter 比较 LangChain 与 CrewAI 编排及其原生路由的分工
OpenRouter 发文将多模型编排拆为三层:工作流编排(LangGraph、CrewAI 负责)、模型路由和提供商路由(OpenRouter 负责)。
Tianyi Cui@tianyi精选AI 评分6565Claude Code 团队发布 Mods,用户通过提示词即可自定义 Claude 的行为和外观,并可将 mods 作为插件分享。
引用Boris Cherny@bchernyMods are absolutely insane. You can now customize Claude to work and look the way you want by just prompting it. Each person works differently, so there's no reason why everyone should have an identical Claude experience. Make Claude your own, and share mods as plugins so others can try your mods too.
推荐理由:作者借 Claude Code Mods 之机,指出 DeepSeekHarness 从一开始就把模型、工具、UI 等全部做成可替换插件,可通过提示词在线修改并持久化。
Andrew Milich@milichabAI 评分4949在 /dashboard 视图中同时管理多个智能体和 worktree
引用Grok@grokNow in Grok Build: A new Agent Dashboard. All your agents on one screen. Try it with /dashboard
AWS Machine Learning BlogAI 评分3939 在 Amazon Bedrock AgentCore 上用智能体 AI 扩展云迁移
AWS Professional Services 构建的四智能体模式在 Amazon Bedrock AgentCore 上运行,将每个应用的基础设施即代码(IaC)开发时间从 3 到 4 周缩短至几分钟,该模式已在一个覆盖 300+ 应用的迁移项目中验证。
Hacker News popular via buzzing.ccAI 评分6060 Earendil 发布 Pi 1.0 并推出实验性持久化智能体框架 Pi Durable
Earendil 与 Pi 社区发布 Pi 1.0,同时推出实验性新包 Pi Durable,一个面向长时运行、持久化、可随处运行智能体的 harness。
Charlie Holtz@charlieholtzAI 评分5959Conductor 发布移动端 Conductor Mobile,用户可以在 iPhone 上运行一支云端智能体团队,现已上线 App Store。

Dongxi 东锡 NLP@dongxi_nlpAI 评分4747引用Tavus@tavusIntroducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).
Replit ⠕@ReplitAI 评分2020
dex@dexhorthyAI 评分2323靠,我知道我一直没做那个"9 个电子游戏"的东西,但这个应该能给你所有需要的信息了 🤣 用 Opus 5.5 做的,还跟 @ElevenLabs 折腾了好久
引用humanlayer@humanlayer_devbeen a little quiet lately but excited to share what's next: HumanLayer is building the software forge for AI and whatever comes after it
Alexandr Wang@alexandr_wangAI 评分3434muse 能监控取消情况,帮你拿到最后一刻极难约到的预约和订位!
引用Dan Peguine@danpeguineMuse just booked for me an appointment with a very busy doctor that had no openings until June, for today. Yesterday I asked it to put a watch on the queue and book as soon as there’s a cancellation. Amazing.
Boris Cherny@bcherny精选AI 评分7070引用ClaudeDevs@ClaudeDevsYou can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
推荐理由:原文介绍了用提示词自定义 Claude Code 行为和界面的新机制,以及通过插件安装和分享的方式。
Artificial Analysis@ArtificialAnlysAI 评分4040Artificial Analysis 的 Coding Agent Index 新增安全拒绝报告,可查看拒绝发生在任务提示词阶段还是智能体已开始工作之后,以及拒绝后切换到了哪个回退模型。

Thariq@trq212AI 评分1616我通常为开发者写作,但下一篇帖子我打算面向那些试图应对 AI 编程智能体变革的公司领导者。 你希望你的领导层能理解智能体的哪些方面?或者如果你正在经营一家公司——你遇到了什么问题?
Andrew Milich@milichabAI 评分3535如果你在使用 Gemini Enterprise Agent Platform,Grok 4.7 现已可用
引用SpaceXAI@SpaceXAIGrok 4.7 is now available on the Gemini Enterprise Agent Platform
SpaceXAI@SpaceXAIAI 评分2828Grok 4.7 现已在 Gemini Enterprise Agent Platform 上线

Karina@karinanguyenAI 评分3434
引用Thoughtful@thoughtfullabPostTrainBench v1.2 is out! A few updates: 1. Cloud GPU support. You can now run the benchmark with identical settings through Harbor + Modal using our new Harbor adapter. 2. New leaderboard leaders. Fable 5.1 takes #1 at 44.6%, followed by Opus 5.5 at 43.8% and GPT-6 (Astra) at 41.9%. 3. Evaluation fixes. Removed BFCL, fixed HumanEval and remote-code scoring, added averaging across multiple seeds, and switched contamination checks to majority vote.
Michael Truell@mntruellAI 评分2828引用Grok Bot@botGrok Bot can now suggest ways to help without you needing to ask.
Thariq@trq212AI 评分6363引用ClaudeDevs@ClaudeDevsYou can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
Anthropic@AnthropicAIAI 评分4242
Claude@claudeaiAI 评分6262引用Claude@claudeaiYou can also now make decks, docs, and designs in your conversation. Draft the one-pager in Claude Docs, turn it into a deck with Claude Slides, and mock up a matching visual in Claude Design, all from one place.
ClaudeDevs@ClaudeDevs精选AI 评分6666
推荐理由:官方给出 mods 的可改动范围和安装方式,读者可以据此判断是否用 TypeScript 定制自己的 Claude Code 工作流。
Elon Musk@elonmuskAI 评分4848引用Beff (e/acc)@beffjezosGrok Bots have been life-changing for someone like me with ADHD who has no patience for context switching / navigating slow interfaces to retrieve information We're seeing the beginnings of personal superintelligence that augments each humans to realize their full potential
Odyssey@odysseymlAI 评分2828今天我们推出 PROWL-2,智能体及其世界模型通过递归学习实现改进。 智能体暴露想象中的错误,修复这些错误又能促成进一步的学习。 我们相信,这种开放式学习是迈向超级智能的关键一步。

AWS Machine Learning BlogAI 评分5252 AWS 博客演示如何用 NVIDIA NeMo Agent Toolkit 和 Amazon S3 Vectors 构建智能体记忆
AWS 博客发布教程,演示如何将 Amazon S3 Vectors 作为 NVIDIA NeMo Agent Toolkit(NAT 1.6)的自定义记忆后端,并部署在 Amazon EKS 上。
TechCrunch · AIAI 评分5656 AWS 发布开源决策模型 Strands Decider 2B,基于 Qwen3.5-2B
AWS 发布开源决策模型 Strands Decider 2B,灵感来自 TypeSafe 的 Jev,可在预设选项间高速低成本地做选择并给出置信度。模型完全开源、可本地运行,由 Amazon 杰出工程师 Marc Brooker 的内部项目改进而来,基于 Qwen3.5-2B 的架构但不生成文本,而是输出校准后的选择,同一周 OpenAI 也宣布了类似产品。
TechCrunch · AIAI 评分5757 Shopify 发布 Canvas,商家可通过与 AI 对话搭建在线商店
Shopify 推出建站工具 Canvas,商家通过与 Shopify 的 AI 智能体 Sidekick 对话即可搭建商店。Canvas 实时渲染店铺真实代码而非静态预览,可测试页面交互、动画并查看不同屏幕尺寸下的效果;初始版本暂不支持第三方主题、app blocks、翻译等,且仅限桌面端。
AWS Machine Learning BlogAI 评分5151 AWS 用 Amazon Bedrock AgentCore 构建环境智能体:从事件驱动信号到人机协同工作流
AWS 发布基于 Amazon Bedrock AgentCore 的环境智能体(ambient agent)参考实现,用 S3 上传等事件自动触发 agent 任务,替代等待用户输入的聊天模式。
elvis@omarsar0AI 评分5757引用Mac Liu@themacliuI’m excited to announce that @arceuslegal is launching with $17M in funding, led by @greycroftvc, with participation from @craft_ventures, @spc, and others. As a founder, I always hated how helpless I felt working with law firms. I went through four or five different firms and somehow the experience was always the same. I’d be waiting on something important to our business with no idea when I’d hear back. I’d have to re-explain our business over and over again. And I dreaded jumping on calls because I knew every minute was costing me money. We started Arceus because we believe every business deserves a better law firm. One that moves faster, costs less, and puts the client first. And we’re just getting started. ↓
赵纯想@chunxiangaiAI 评分1919邮件均已收到,节后回复。 我 10 月下旬到杭州,先找个办公室。 然后我们开干。 前 50 名员工都有财富自由的风险。
引用赵纯想@chunxiangai正式宣告:如果你想在今年冬天,在 *杭州* 大展身手。把微信、telegram,重新写一次。以 AI Native 的方式,去构建一个 Agent 与人,共为一等公民的 IM 网络。请您提前与我联系。 一直以来,Cromma(可爱信)无法真正通过思想实验。所有的壁垒都有解法,可优化的点千千万万,唯独小程序生态这一块,完全难以撼动。 直到这个秋天,我们看到了很多东西真正进入了沸腾阶段(以lovable为标志的产品)。让每个用户用自然语言的方式,0代码、0部署焦虑地来创造和分发自己的“小程序”。将IM中的可流转软件生态,从少部分人编码,大部分人使用的时代,蝶变到程序的创造者、迭代者、使用者,都是用户本人的时代。 想象一下:hi,为我的这个学员群弄一个小程序。 总而言之,一切已经启动。我需要做一些 CEO 该做的事。组织人,尤其是野心勃勃的人。我们一起打造,一个唯一的 Rust IM 内核。经 UniFFI 绑定给各端原生。经 wasm 给 web,经 napi 给 Electron。 这里放不下关于Cromma的想象。让一切从一封邮件开始,只需证明你也想干,且可以胜任。 cromma@laper.ai
Aravind Srinivas@AravSrinivasAI 评分5252引用Perplexity@perplexity_aiPerplexity Computer now creates interactive charts and visualizations directly in your thread. For financial data, Computer uses @tradingview Lightweight Charts for candlesticks, volume, and moving averages.
elvis@omarsar0AI 评分5858Meta 及合作机构发布 Context Language Models(CLMs)论文,把上下文作为文件让模型用 Bash 自由读写编辑,自行决定保留、重写或删除内容。
