跳到正文

#Agent

今日 81 条
今天10月2日周五
  1. elvis48

    AgentWorld 将 3 到 20 个不同角色的 LLM 智能体放入游戏沙盒,执行 50+ 轮的长程任务,智能体无法看到彼此内部状态,只能通过消息和共享计划协调。Gemini 3 Flash 任务成功率最高,为 52.0%;协调类任务最难,成功率仅 12%,常见失败包括沟通中断、角色混淆和共享计划丢失。论文显示,多智能体团队中不到三分之一的行动真正有助于完成任务。

  2. Hugging Face Daily Papers47

    AutoGUIWorld:用图像生成器作为 GUI 智能体的视觉世界模型

    AutoGUIWorld 是一个数据生成框架,结合图像生成器的视觉先验与规划器的任务知识,无需部署或运行软件环境即可合成 GUI 交互轨迹。它从操作系统上下文、视觉外观和界面状态的结构化规格中采样初始场景,生成 79,266 条覆盖 Ubuntu、Windows、macOS 和 Chrome 的空间标注步级训练样本。

  3. elvis50

    我认为更快的推理是编程智能体下一个重大突破之一。 Volantis 正在利用光学技术为每颗芯片提供远超以往的内存和更高的内存带宽。他们的目标是在超过 10T 参数的模型上,实现每用户每秒高达 10,000 tokens。这太疯狂了! 以那种速度,今天需要数小时的编程智能体可以在几分钟内完成。 绝对是我最近见过的最令人兴奋的融资之一。

    引用Tapa Ghosh@semiDL

    Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post

  4. OpenRouter Announcements64

    OpenRouter 详解六大 Agent 框架的工具调用 Schema 处理,并提出在 API 层统一格式

    OpenRouter 比较了 LangChain、LangGraph、CrewAI、OpenAI Agents SDK、Claude Agent SDK、Microsoft Agent Framework 和 Google ADK 如何定义工具 Schema 并在不同提供商的 wire format 之间做翻译。

    推荐理由:原文逐一拆解六大框架的工具调用格式翻译位置,并给出在 API 层统一格式的可行做法,便于开发者选型前对齐自己的技术栈。

  5. Tianyi Cui65

    Claude Code 团队发布 Mods,用户通过提示词即可自定义 Claude 的行为和外观,并可将 mods 作为插件分享。

    引用Boris Cherny@bcherny

    Mods are absolutely insane. You can now customize Claude to work and look the way you want by just prompting it. Each person works differently, so there's no reason why everyone should have an identical Claude experience. Make Claude your own, and share mods as plugins so others can try your mods too.

    推荐理由:作者借 Claude Code Mods 之机,指出 DeepSeekHarness 从一开始就把模型、工具、UI 等全部做成可替换插件,可通过提示词在线修改并持久化。

  6. Dongxi 东锡 NLP47

    Tavus 推出 Griffin,号称首个通过视频图灵测试的模型,48% 的实时对话者认为它是真人,此前系统通过率不足 3%,并在 NVIDIA 全双工 AI 视频基准上排名第一。它是首个 Human Interaction Model(HIM)。推文作者借此调侃:用 Agent 干活、Griffin 开会甚至面试,就能同时接成百上千个远程职位。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

  7. Boris Cherny70

    Claude Code 推出 Mods 功能,可通过提示词自定义 Claude 的行为、UI 和功能。Mods 用几行 TypeScript 编写或由 Claude 生成,随 plugins 分发,可在 CLI 或桌面应用中用 /plugin 安装,也可作为插件分享给他人使用。

    引用ClaudeDevs@ClaudeDevs

    You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:

    推荐理由:原文介绍了用提示词自定义 Claude Code 行为和界面的新机制,以及通过插件安装和分享的方式。

  8. Karina34

    一个有趣的结果:Fable 5.1 得分 44.6%,高于 Fable 5 的 37.3%,同时平均少用 33 分钟。这里的进步也意味着从每小时的自主工作中获得更多产出。 大多数智能体运行 8–10 小时,却取得显著不同的分数。下一个有趣的问题是,每多一个小时,每个智能体能获得多少提升!

    引用Thoughtful@thoughtfullab

    PostTrainBench v1.2 is out! A few updates: 1. Cloud GPU support. You can now run the benchmark with identical settings through Harbor + Modal using our new Harbor adapter. 2. New leaderboard leaders. Fable 5.1 takes #1 at 44.6%, followed by Opus 5.5 at 43.8% and GPT-6 (Astra) at 41.9%. 3. Evaluation fixes. Removed BFCL, fixed HumanEval and remote-code scoring, added averaging across multiple seeds, and switched contamination checks to majority vote.

  9. Thariq63

    Claude Code 推出 mods 功能,用户可以改变其行为、自定义 UI 并替换自己的功能。mod 可用几行 TypeScript 编写或由 Claude 生成,随插件分发,通过 /plugin 在 CLI 或桌面应用安装。作者表示软件正变得可塑,希望更多软件具备这种可扩展性。

    引用ClaudeDevs@ClaudeDevs

    You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:

  10. Claude62

    Claude 宣布为期两周的活动,在 Claude 应用中以设计、幻灯片或文档开启对话后,该会话的后续工作消耗用量额度减少 50%。活动推荐使用 Claude Sonnet 5.5,并可在同一对话中用 Claude Docs 起草文档、Claude Slides 做幻灯片、Claude Design 做配图。

    引用Claude@claudeai

    You can also now make decks, docs, and designs in your conversation. Draft the one-pager in Claude Docs, turn it into a deck with Claude Slides, and mock up a matching visual in Claude Design, all from one place.

  11. Elon Musk48

    试试 Grok @Bot!

    引用Beff (e/acc)@beffjezos

    Grok Bots have been life-changing for someone like me with ADHD who has no patience for context switching / navigating slow interfaces to retrieve information We're seeing the beginnings of personal superintelligence that augments each humans to realize their full potential

  12. TechCrunch · AI56

    AWS 发布开源决策模型 Strands Decider 2B,基于 Qwen3.5-2B

    AWS 发布开源决策模型 Strands Decider 2B,灵感来自 TypeSafe 的 Jev,可在预设选项间高速低成本地做选择并给出置信度。模型完全开源、可本地运行,由 Amazon 杰出工程师 Marc Brooker 的内部项目改进而来,基于 Qwen3.5-2B 的架构但不生成文本,而是输出校准后的选择,同一周 OpenAI 也宣布了类似产品。

  13. elvis57

    Elvis Saravia 转发 Arceus Legal 融资消息并提出观点,预计更多 AI 公司会自己运营服务并出售结果。Arceus Legal 宣布获得 1700 万美元融资,由 greycroftvc 领投;客户通过 Slack 发送合同,AI 收集上下文,执业律师审查每份工作产出,合同审查平均 3 到 5 小时完成,按开工前约定的固定费用收费。

    引用Mac Liu@themacliu

    I’m excited to announce that @arceuslegal is launching with $17M in funding, led by @greycroftvc, with participation from @craft_ventures, @spc, and others. As a founder, I always hated how helpless I felt working with law firms. I went through four or five different firms and somehow the experience was always the same. I’d be waiting on something important to our business with no idea when I’d hear back. I’d have to re-explain our business over and over again. And I dreaded jumping on calls because I knew every minute was costing me money. We started Arceus because we believe every business deserves a better law firm. One that moves faster, costs less, and puts the client first. And we’re just getting started. ↓

  14. 赵纯想19

    邮件均已收到,节后回复。 我 10 月下旬到杭州,先找个办公室。 然后我们开干。 前 50 名员工都有财富自由的风险。

    引用赵纯想@chunxiangai

    正式宣告:如果你想在今年冬天,在 *杭州* 大展身手。把微信、telegram,重新写一次。以 AI Native 的方式,去构建一个 Agent 与人,共为一等公民的 IM 网络。请您提前与我联系。 一直以来,Cromma(可爱信)无法真正通过思想实验。所有的壁垒都有解法,可优化的点千千万万,唯独小程序生态这一块,完全难以撼动。 直到这个秋天,我们看到了很多东西真正进入了沸腾阶段(以lovable为标志的产品)。让每个用户用自然语言的方式,0代码、0部署焦虑地来创造和分发自己的“小程序”。将IM中的可流转软件生态,从少部分人编码,大部分人使用的时代,蝶变到程序的创造者、迭代者、使用者,都是用户本人的时代。 想象一下:hi,为我的这个学员群弄一个小程序。 总而言之,一切已经启动。我需要做一些 CEO 该做的事。组织人,尤其是野心勃勃的人。我们一起打造,一个唯一的 Rust IM 内核。经 UniFFI 绑定给各端原生。经 wasm 给 web,经 napi 给 Electron。 这里放不下关于Cromma的想象。让一切从一封邮件开始,只需证明你也想干,且可以胜任。 cromma@laper.ai

  15. Aravind Srinivas52

    Perplexity 开始在 Computer 中推出内嵌图表、图形和可视化功能。金融数据方面,Computer 使用 TradingView 的 Lightweight Charts 提供 K 线、成交量和移动平均线,图表在会话线程中直接生成且可交互。

    引用Perplexity@perplexity_ai

    Perplexity Computer now creates interactive charts and visualizations directly in your thread. For financial data, Computer uses @tradingview Lightweight Charts for candlesticks, volume, and moving averages.