跳到正文

全部动态

今日 75 条
9月26日周六
  1. ClaudeDevs70

    Opus 5.5 的输入和输出 token 比 Opus 5 便宜 20%,cache reads 便宜 60%。作者据此测算了在 Claude Code 中完成一个任务的实际成本变化,并发布了博客和计算器,读者可从 /usage 运行自己的数据:https://claude.dev/blog/what-a-task-costs-on-opus-5-5/

    推荐理由:作者用具体数字拆解了 Opus 5.5 降价对 Claude Code 单任务成本的实际影响,并给出可复用的成本计算器入口。

  2. GitHub Blog · AI & ML60

    GitHub Copilot app 新手教程:如何用 canvases 构建自定义工作流

    GitHub Copilot app 提供 canvas 功能,运行 /create-canvas 技能并用自然语言描述工作流、界面操作和 agent 职责三要素,即可生成看板、清单等自定义界面并保存为可共享的 extension。

    推荐理由:官方教程给出从描述到生成可复用 canvas 界面的具体步骤和三个提问框架,读者可以直接照做迁移到自己的工作流。

  3. Anthropic65

    Anthropic 发文称,Claude 在收到单个九圈问题提示词后,在 Claude Science 中基本无人监督地运行数天,用 Dixon 等人的方法完成求解,总成本几千美元,突破了此前八圈的纪录(平面 N=4 超杨-米尔斯简化模型)。物理学家 Lance Dixon 独立验证了结果,von Hippel 为该博客撰写了经历回顾。

    推荐理由:九圈散射振幅计算由 Dixon 独立验证,计算成本仅几千美元,为学界评估 AI 科研能力提供了一个可核验的案例。

  4. Boris Cherny50

    Boris Cherny 称 Claude Tag 每天写他超 50% 的 PR,完成约 100% 的数据分析,并修复大部分产品反馈和 bug。他介绍 Claude Tag 不同于普通 Slack bot,具备主动、可编程、有记忆和连接器访问能力,配合 Opus 5.5 和 Fable 5.1 有较强判断力,并给出自动复现 bug 并提 PR、深挖数据假设、生成讲解游戏等示例提示词。

    引用Noah Zweben@noahzweben

    Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels

  5. Noah Zweben49

    今年 2 月我们首次推出 /remote-control 时,我用 Opus 4.6 做这些视频玩得很开心。那么,这是 Opus 5.5 的粘土动画版本。 Remote Control with 5.5,当你不得不去的时候!

    引用Noah Zweben@noahzweben

    Rolling out Claude Code Remote Control to Pro users - because they deserve to use the bathroom too . (Team and Enterprise coming soon). 🧻 Rolling out to 10% and ramping 1. Update to claude v2.1.58+ 2. Try log-out and log-in to get fresh flag values. 3. /remote-control

  6. Claude66

    Claude 官方表示 Claude Opus 5.5 发布数日,汇总了用户用其探索和发现的喜爱案例。引用案例中,@RyanSael 让 Opus 5.5 通过构建交互式镜头实验室讲解相机对焦,一次生成耗时 1 小时 26 分钟,API 成本 $25.66,成品见 https://lens.lab.sael.net。

    引用Ryan Sael@RyanSael

    I asked Opus 5.5 to explain camera focus by building an interactive lens lab Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost https://lens.lab.sael.net Move the focus ring and you can see the glass elements shift the sharp plane through the scene

    推荐理由:官方汇总用户用 Opus 5.5 探索的成果,引用案例给出了单次生成时长与成本,可作实际使用参考。

9月25日周五
  1. karminski-牙医59

    美团 LongCat 发布 LongCat-2.5-Preview 模型,总参数 1.6T、激活约 48B,支持 1M token 上下文窗口,原生多模态,面向终端、浏览器、GUI、表格和设计工具等长程任务。API 与聊天入口分别见 https://longcat.ai/platform/ 和 https://longcat.ai/chat/;作者补充定价与之前一样,图中显示输入(缓存未命中)2.00 元/百万 token,输入(缓存命中)0.04 元/百万 token,输出 8.00 元/百万 token。

    引用Meituan LongCat@Meituan_LongCat

    LongCat-2.5-Preview is now live. 1.6T parameters. ~48B active. A 1M-token context window. Natively multimodal. Built to take on long-horizon tasks. From terminals and browsers to GUIs, spreadsheets, and design tools. Try it now: 🚀 API: https://longcat.ai/platform/ 💬 Chat: https://longcat.ai/chat/

  2. swyx24

    今年 1 月,我给自己的内容策略定了个目标:“Scaling without Slop”。 它终于开始奏效了。我们花了 3 年才在 YouTube 上达到第一个 10 万订阅。而接下来的 10 万只用了 1.2 个月。AEO/SEO/订阅增长等其他指标也类似,我还有很多 New Media 的想法,很期待去实现。 正式预告 Latent Space、AINews 以及 swyx inc 其他业务下一阶段的发展,详见下方

    引用Latent.Space@latentspacepod

    [AINews] The Future of Latent Space https://www.latent.space/p/ainews-the-future-of-latent-space - Plans for AINews v3 - Plans for a new home! - We are open for business - and @supabase are our first sponsors!

  3. François Chollet43

    我认为软件工程的"难度"本质上是恒定的,无论你迁移到哪个抽象层级,因为人类认知会适应新工具,直到能够充分发挥自身能力。 工具只是可供性,不是让工作消失的魔法棒。 伟大的软件工程以前极其困难。现在依然极其困难,尽管工作流程已大不相同。

    引用Simon Willison@simonw

    The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge

  4. Hugging Face Daily Papers40

    DMM:通过迭代意图去噪实现多智能体路径规划的联合动作优化

    DMM(Decentralized Master-Mind)提出用离散的迭代意图精炼替代一次性动作采样,解决多智能体路径规划中局部有效动作重组为不兼容联合动作的问题。该方法受扩散模型去噪启发,通过通信轮次耦合智能体选择,并用模仿学习预训练、MICPO 强化学习优化。在 1,600 个 MovingAI 任务上,DMM 经 MICPO 微调后解决 1,598 个,覆盖率最高,且可扩展至超百万智能体。

  5. Hugging Face Daily Papers35

    RWTD:用奖励加权传输蒸馏对齐一步生成模型

    研究提出奖励加权传输蒸馏(RWTD),一种仅需生成样本和标量奖励评估的一步生成模型后训练方法,通过混合当前与参考分布的自适应目标,经特征空间最优传输与不动点回归实现。该方法将一步 SANA Sprint 1.6B 的 GenEval 分数从 0.73 提升至 0.80,偏好对齐实验显示其具备跨奖励泛化能力,能兼顾改进与组合能力保留。

  6. GitHub Blog · AI & ML61

    GitHub Copilot 博客:为什么聊天界面往往不是正确的 UI,用 canvas 试试

    GitHub Copilot 博客作者提出聊天(chat)很多时候是错误的 AI 交互界面,介绍 GitHub Copilot app 中的 canvas,它是在应用内运行、无浏览器外壳的全栈小应用,可与 Copilot agent 双向通信。

    推荐理由:作者作为 Copilot 团队成员提出用 canvas 自定义界面替代聊天框,并结合工作流自动化等实例说明何时值得让智能体先造工具。