跳到正文

#编码

今日 26 条
9月27日周日
  1. Thomas Wolf27

    我们曾有过一段亲手雕琢代码的美好时光,如今它结束了。 在另一面,是一段激动人心的全新职业生涯——成为专业的造物者,驾驭那些直到不久前还只存在于科幻中的智能。 能亲身经历那个一切靠双手完成的时代,又恰好站在切换的那一刻,何其有幸。

    引用Scott@scottstts

    My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do

9月26日周六
  1. Hugging Face Daily Papers43

    CUA-SWE:当 Computer-Use Agent 遇上可视化软件工程

    研究者推出 CUA-SWE,一个面向"计算机使用+软件工程"的 benchmark、环境与评测流水线,覆盖四个软件工程领域,要求智能体在同一任务内改代码与配置、执行命令、操作运行中的软件并查看视觉反馈。每个任务配有确定性的专属测试,用于验证软件是否满足需求并保持既定行为。该工作评测前沿智能体如何结合源码级执行与应用截图、图形交互,产出经过验证的软件改动。

  2. ClaudeDevs70

    Opus 5.5 的输入和输出 token 比 Opus 5 便宜 20%,cache reads 便宜 60%。作者据此测算了在 Claude Code 中完成一个任务的实际成本变化,并发布了博客和计算器,读者可从 /usage 运行自己的数据:https://claude.dev/blog/what-a-task-costs-on-opus-5-5/

    推荐理由:作者用具体数字拆解了 Opus 5.5 降价对 Claude Code 单任务成本的实际影响,并给出可复用的成本计算器入口。

  3. Boris Cherny50

    Boris Cherny 称 Claude Tag 每天写他超 50% 的 PR,完成约 100% 的数据分析,并修复大部分产品反馈和 bug。他介绍 Claude Tag 不同于普通 Slack bot,具备主动、可编程、有记忆和连接器访问能力,配合 Opus 5.5 和 Fable 5.1 有较强判断力,并给出自动复现 bug 并提 PR、深挖数据假设、生成讲解游戏等示例提示词。

    引用Noah Zweben@noahzweben

    Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels

  4. Noah Zweben49

    今年 2 月我们首次推出 /remote-control 时,我用 Opus 4.6 做这些视频玩得很开心。那么,这是 Opus 5.5 的粘土动画版本。 Remote Control with 5.5,当你不得不去的时候!

    引用Noah Zweben@noahzweben

    Rolling out Claude Code Remote Control to Pro users - because they deserve to use the bathroom too . (Team and Enterprise coming soon). 🧻 Rolling out to 10% and ramping 1. Update to claude v2.1.58+ 2. Try log-out and log-in to get fresh flag values. 3. /remote-control

  5. Claude66

    Claude 官方表示 Claude Opus 5.5 发布数日,汇总了用户用其探索和发现的喜爱案例。引用案例中,@RyanSael 让 Opus 5.5 通过构建交互式镜头实验室讲解相机对焦,一次生成耗时 1 小时 26 分钟,API 成本 $25.66,成品见 https://lens.lab.sael.net。

    引用Ryan Sael@RyanSael

    I asked Opus 5.5 to explain camera focus by building an interactive lens lab Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost https://lens.lab.sael.net Move the focus ring and you can see the glass elements shift the sharp plane through the scene

    推荐理由:官方汇总用户用 Opus 5.5 探索的成果,引用案例给出了单次生成时长与成本,可作实际使用参考。

9月25日周五
  1. François Chollet43

    我认为软件工程的"难度"本质上是恒定的,无论你迁移到哪个抽象层级,因为人类认知会适应新工具,直到能够充分发挥自身能力。 工具只是可供性,不是让工作消失的魔法棒。 伟大的软件工程以前极其困难。现在依然极其困难,尽管工作流程已大不相同。

    引用Simon Willison@simonw

    The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge

  2. GitHub Blog · AI & ML61

    GitHub Copilot 博客:为什么聊天界面往往不是正确的 UI,用 canvas 试试

    GitHub Copilot 博客作者提出聊天(chat)很多时候是错误的 AI 交互界面,介绍 GitHub Copilot app 中的 canvas,它是在应用内运行、无浏览器外壳的全栈小应用,可与 Copilot agent 双向通信。

    推荐理由:作者作为 Copilot 团队成员提出用 canvas 自定义界面替代聊天框,并结合工作流自动化等实例说明何时值得让智能体先造工具。

  3. Claude Code GitHub Releases35

    Claude Code v2.1.282 发布

    Claude Code 发布 v2.1.282,新增 maxProseWidth 设置,可限制宽终端中 Claude 正文宽度,表格与代码块仍保持全宽。该版本还新增启动提示及 /status、claude doctor 条目,列出项目设置文件中被忽略或关闭遥测的变量,并修复了会话续接重发旧消息、扩展思考丢失及多处 API 报错等问题。

  4. GitHub Blog · AI & ML63

    GitHub Security Lab 发布 Fuzzing Taskflow:用 LLM 智能体自动化 C/C++ 模糊测试

    GitHub Security Lab 的 Antonio Morales 基于自家的 Taskflow Agent 框架构建了 Fuzzing Taskflow,一个面向 C/C++ 项目的自主模糊测试流水线。

    推荐理由:原文给出完整的架构设计、覆盖反馈循环和分层判断方法,读者可以据此把 LLM 智能体接到自己的模糊测试流程里。

9月24日周四
  1. inclusionAI Hugging Face models62

    inclusionAI 开源 Ling-flash-2.0:100B 总参数、6.1B 激活的 MoE 模型

    inclusionAI 正式开源 Ling 2.0 架构下的第三个 MoE 大语言模型 Ling-flash-2.0,总参数 100B、激活参数 6.1B(非嵌入 4.8B),基于 20T+ tokens 数据训练并经 SFT 和多阶段强化学习。

    推荐理由:官方给出参数结构、基准对比和推理速度数据,读者可据此评估小激活 MoE 替代 40B 稠密模型的可行性。

  2. Claude Code GitHub Releases39

    Claude Code v2.1.281 发布

    Claude Code v2.1.281 为 Claude apps gateway 新增 Bedrock upstream 的 assume_role 与 guardrail 配置,并支持在 settings.json 中设置 "attribution": false 隐藏提交与 PR 署名。该版本还修复了会话恢复时重发历史、提示缓存丢失、代理中断流被误判为完成等多项问题。

9月23日周三
  1. Latent Space85

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日跟进低价 GPT-6 Sol 与 Luna

    AINews 汇总 Claude Opus 5.5 发布:作为 Claude 5.5 家族首个模型,官方称多数任务达到 Fable 5.1 水平、比 Opus 5 快约 30% 且每任务便宜约 40%,token 单价从 $5/$25 降至 $4/$20,并成为 Claude Code 与 Claude 应用的新默认模型。

    推荐理由:汇总了 Opus 5.5 与 GPT-6 Sol/Luna 同日发布的基准、定价与第三方评测,含实际每任务成本与争议细节,便于对比两家宣传口径。

  2. Google Developers Blog65

    Antigravity SDK 支持本地模型,首发接入 Gemma 4 26B A4B 与 LiteRT

    Google 宣布 Antigravity SDK 支持本地工作流,首发支持通过 Google AI Edge 的 LiteRT 运行 Gemma 4 26B A4B,可完全离线提供智能体能力,建议机器配备 24GB 以上 VRAM 或统一内存。

    推荐理由:原文给出本地运行智能体的具体配置方法和混合编排演示数据,开发者可以直接照此把智能体工作流搬到本地 GPU 上。

  3. Simon Willison83

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 与 GPT-6 Luna 并掀起价格战

    Anthropic 于 9 月 22 日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 定价 $0.10/$0.50 每百万 token,比 GPT-5.6 Luna 再降一半;Opus 5.5 降价 20% 至 $4/$20,缓存读取价格下降 60%。

    推荐理由:作者用实测和价格对比表梳理了这轮降价的具体幅度,还发现 Opus 5.5 max 档过度思考撞上输出上限的问题。

  4. Greg Brockman79

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna 两款更快、更实惠的模型,基于 GPT-6 Astra 的技术积累。两款模型在专业工作、事实性、编码、computer use 和对齐方面延续 Astra 的 SOTA 表现,同时缓存和推理效率提升使 API 价格比 GPT-5.6 促销价低 50%。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:原文由当事方宣布两款新模型及降价幅度,读者可以据此了解 GPT-6 系列的能力分工与成本变化。

  5. Boris Cherny76

    Boris Cherny 称 Claude Opus 5.5 是他最近几周的日常主力模型。他让 Opus 5.5 和 Fable 5.1 各把 HAProxy 从 C 移植到 Rust,两者都几乎通过全部测试,但 Opus 5.5 用时 9.5 小时,Fable 5.1 用时 12 小时,且成本低 51%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:作者亲测对比两个模型移植 HAProxy 的耗时与成本,给出了具体数字供选型参考。

  6. Claude Code GitHub Releases74

    Claude Code v2.1.280 发布,Claude Opus 5.5 成为默认 Opus 模型

    Claude Code v2.1.280 加入 Claude Opus 5.5(claude-opus-5-5)作为默认 Opus 模型,1M 上下文,$4/$20 per Mtok,cache reads $0.20/Mtok;Pro 和 Team Standard 计划默认模型由 Sonnet 改为 Opus。

    推荐理由:发布日志列出 Opus 5.5 成为默认模型及其定价上下文,并包含大量修复与 VSCode 新对话框,读者可对照更新自己的工作流。

  7. Claude Blog67

    Anthropic 工程师分享如何为 AI 驱动的代码现代化项目做准备

    Anthropic 前线部署工程师在 Notes from the Field 系列中分享管理大规模代码现代化项目的经验:原本需数年的项目现在可在数月或数周内完成,但瓶颈从编写变更转移到组织动员。

    推荐理由:来自 Anthropic 一线工程师的实战经验,给出 AI 代码现代化项目启动前的组织性准备工作清单和可复用六步流程。

9月22日周二
  1. elsewhere articles65

    Step 5 Preview 实测:榜单之外的真实表现与短板

    阶跃毫无预兆发布 Step 5 Preview,总参数量 600B、激活 27B,带视觉输入,称 Artificial Analysis 上涨 44 分、全球开源前三,单任务成本仅为 Claude Opus 5 的 1/8。

    推荐理由:作者与朋友实测了 Step 5 Preview 在可视化、金融、游戏等任务上的真实表现,补上了榜单分数之外的第一手体感参考。