跳到正文

AI 编码

AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。

当前仅显示精选新闻
436条精选相关主题Agent 智能体Cursor教程实践

最新精选

第 321–340 条 · 共 436 条
6月9日周二
  1. @berryxia67

    Kimi Code 开源编码智能体迎来重大升级,支持一行命令 CLI 安装、零配置快速启动,并可拖入视频作为编码上下文,实现参考视频转 LUT、长视频转短视频、屏幕录像转代码。插件系统可拉取股票价格、财报和学术论文,同时支持 ACP 协议接入 JetBrains 和 Zed,并提供自定义 hooks 扩展工作流,可搭配 Kimi K2.6 使用。

    引用Kimi Developers (@KimiDevs)@KimiDevs

    Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-recording-to-code, and more​ 🔹Plugins for stocks, financial reports, academic papers, with more coming​ 🔹Supports the ACP protocol, and works with JetBrains, Zed, and more​ 🔹Hooks for custom tools and workflows​ Try it with Kimi K2.6 👉 kimi.com/code Issues, plugin ideas, and PRs welcome! Community feedback helps shape what ships next.​🚀

    推荐理由:原文列出升级后的零配置安装、视频上下文与插件能力,读者可据此判断编码智能体的门槛变化。

6月8日周一
  1. 量子位 · 微信公众号76

    OpenAI 高管称「Chat 已死」,ChatGPT 将改版为 Agent 超级应用

    OpenAI 高管提出「Chat 已死」,公司正推进 ChatGPT 诞生以来最大规模改版,目标是从聊天机器人变成个人 Agent 式的超级应用。改版由 Codex 承担 Agent 能力,其周活已超 500 万、非开发者用户占 20%,并新增可直接操作电脑、并行运行多个 Agent 等能力。动因是 ChatGPT 虽在 5 月突破 10 亿月活但多数用户免费,企业端支出正被 Claude 抢占。

    推荐理由:文章梳理了 OpenAI 把 ChatGPT 从聊天框转向 Agent 超级应用的路线,并用 Codex 与 Claude 的数据呈现其转型压力。

6月5日周五
  1. 数字生命卡兹克 · 微信公众号79

    Anthropic 发布万字长文《当 AI 开始构建自己》,讨论递归自我改进

    Anthropic 研究院发布长文《当 AI 开始构建自己》,用公开基准和此前未披露的内部数据说明 AI 已在加速 AI 系统自身的开发。文中称 Anthropic 工程师平均每季度交付的代码量是 2021 至 2025 年间的 8 倍,Claude 在最开放任务上的成功率在 2026 年 5 月达到 76%,六个月内提高 50 个百分点。

    推荐理由:Anthropic 用内部数据展示 Claude 在写代码和做研究上的进展,读者可据此理解递归自我改进这一趋势的早期证据。

  2. @kimmonismus66

    Anthropic 发布博客探讨递归自我改进,称距离能完全自主设计并构建后继模型的 AI 已不远,但强调这尚未到来、也非必然,只是可能比多数机构预想的更早。文中引用数据称 Anthropic 工程师如今每季度交付代码量约为 2021–2025 年的 8 倍,AI 能可靠完成的任务时长约每 4 个月翻一倍,截至 2026 年 5 月 Claude 撰写了并入其代码库 80%+ 的代码。博客还给出三种未来路径,认为人类设定方向、效率持续复利提升是可能路径,而完全递归自我改进的对齐结果最不确定。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:文中并列了编码速度、任务时长与代码占比等具体数字,可用来观察 AI 自主编码能力的演进节奏。

  3. @kimmonismus68

    Anthropic 发布博客文章讨论递归自我改进(RSI),称距离能完全自主设计和构建后继模型的 AI 已不远,但强调这尚未实现也并非必然。文中数据包括 Anthropic 工程师如今每季度交付的代码量约为 2021 至 2025 年的 8 倍,截至 2026 年 5 月 Claude 编写了并入其代码库 80% 以上的代码,AI 能可靠完成的任务时长约每 4 个月翻一倍。文章提出三种未来路径,其中人类仍掌握方向的复合效率提升被视为最可能的走向。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:汇总了 Anthropic 博客关于递归自我改进的关键数据与三种未来路径,可据此判断自动化编码的推进速度。

  4. @kimmonismus73

    Anthropic 发布博客称其内部数据显示 Claude 正在加速 AI 研发,存在走向递归自我改进的可能,并强调这尚未到来、也并非必然。博客列举的指标包括:Anthropic 工程师每季度交付代码量约为 2021–2025 年平均水平的 8 倍,AI 能可靠完成的任务时长约每 4 个月翻倍(此前为每 7 个月)。

    引用Anthropic (@AnthropicAI)@AnthropicAI

    Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…

    推荐理由:转述 Anthropic 内部数据,读者可据此了解递归自我改进讨论背后的具体加速指标。

6月4日周四
  1. @RyanLeeMiniMax66

    MiniMax 发布开放权重模型 MiniMax M3,官方称其是首个同时结合编码与智能体、100 万上下文、原生多模态三项前沿能力的开放权重模型。官方给出 59.0% SWE-Bench Pro、66.0% Terminal Bench 2.1 等成绩,通过 MiniMax Sparse Attention 将上下文扩展至 1M,并从 Step Zero 起原生多模态。作者补充 M3 目前位列 ArtificialAnalysis 第 8,因需并行开源 MSA 算子,权重将于下周晚些时候向所有人发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:官方给出三项前沿能力与基准成绩,并说明权重和 MSA 算子的开源时间,读者可据此判断开放节奏与可用范围。

  2. @arena80

    MiniMax M3 登入 Arena,在 Code Arena 前端编码榜排名第 7,得分 1531,与 GLM-5.1 接近。其定价为每 M token 输入 0.60 美元、输出 2.40 美元,Arena 称其在所属价位推动了性价比前沿。MiniMax 官方介绍称 M3 是首个同时具备三项前沿能力的开源权重模型,SWE-Bench Pro 59.0%、Terminal Bench 2.1 66.0%、MCP Atlas 74.2%,通过 Sparse Attention 将上下文扩展至 1M,原生多模态,权重与技术报告约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:MiniMax M3 在 Arena 前端编码榜位列第 7,与 GLM-5.1 接近,读者可据此比较同级价位模型的能力与定价。

  3. @runware65

    MiniMax 发布开源权重模型 M3,称其同时结合编码与智能体、1M 上下文和原生多模态三项能力。官方给出 SWE-Bench Pro 59.0%、Terminal Bench 2.1 66.0%、MCP Atlas 74.2% 等成绩,并称 MiniMax Sparse Attention 将上下文扩展至 1M。Runware 表示已可通过其 API 调用该模型,权重与技术报告约 10 天后发布。

    引用MiniMax (official) (@MiniMax_AI)@MiniMax_AI

    Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: platform.minimax.io Token Plan: platform.minimax.io/subscrib… 🚀New! MiniMax Code: code.minimax.io Weights & Tech Report in ~10 Days

    推荐理由:引用的发布信息列出 M3 在编码、智能体与 1M 上下文上的基准成绩,可供判断这款开源权重模型的能力组合。

6月3日周三
  1. @alexxubyte66

    OpenAI 的数据智能体只用一个模型和 13 个工具,在 1.5 exabytes 数据和 90,000 张表上运行,设计上刻意保持简单。数据平台工程负责人 Emma Tang 介绍了其架构,以及让单一 LLM 在 90,000 张表上保持可靠的六层上下文。文章还包含 OpenAI 内部使用 Codex 的三个场景,以及五条供团队构建领域智能体参考的实践经验。

    推荐理由:多数团队会堆叠路由器、微调与多模型,OpenAI 的数据智能体却只用单模型加 13 个工具跑通 90,000 张表,形成可对比的工程取舍。

  2. @OpenRouter67

    OpenRouter 推出 Pareto Code,一个免费、实验性的编程路由。开发者在请求中设置 min_coding_score,即可路由到满足该门槛的最便宜代码能力模型,排名由 Artificial Analysis 提供,并可实时看到 Pareto 前沿的移动。作者本人的推文仅表示将提供该路由的更多信息。

    引用OpenRouter (@OpenRouter)@OpenRouter

    Introducing Pareto Code: a new, free, experimental coding router Set `min_coding_score` in your request and route to the cheapest code-capable model that clears your bar, ranked by @ArtificialAnlys. See the Pareto frontier shifting in real time👇

    推荐理由:请求可按代码能力门槛自动筛选最便宜可用模型,为控制编码模型成本提供了一种可配置的路由思路。

  3. @berryxia71

    微软 AI 在 Build 上发布七个全新 MAI 模型,官方称并非简单迭代,而是从零开始、干净数据血统、零蒸馏训练的一整个家族,涵盖推理、编码、图像、转录与语音并各有 Flash 版本。

    引用Microsoft AI (@MicrosoftAI)@MicrosoftAI

    Seven new models launching at Build: let’s go! Reasoning. Code. Image. Transcribe. Voice. Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models Thread 🧵 #MSBuild

    推荐理由:文中梳理了七个 MAI 模型的任务分工与基准数字,可据此了解微软从零训练、任务专精的模型家族路线。

  4. Claude Blog63

    Anthropic Claude Code 团队如何重构 AI 原生工程组织的流程

    Anthropic Claude Code 团队负责人复盘了智能体编码成为默认工作方式后对工程流程的重写:规划从六个月路线图改成 just-in-time,遇到问题先问 Claude 并追问能否自动化,代码评审由 Claude 处理风格、缺陷和测试,人只在法律风险、信任边界与安全代码、产品判断等需要领域专长处介入。

    推荐理由:Claude Code 团队把规划、上下文获取、评审与分工按智能体编码重写,并给出可对照的前后变化与跟踪指标。

  5. @OpenAI76

    OpenAI 为 Codex 扩展 plugins,使其不再局限于单个工具。用户一次安装即可让 Codex 成为特定角色的专家,无需编码。Codex 现可访问 62 个热门应用和 110 个技能,覆盖销售、数据分析、创意生产、产品设计和公开股票投资等工作场景。

    推荐理由:原文给出可一次安装的角色化插件、接入应用与技能数量,读者可据此判断 Codex 能覆盖哪些岗位工作。

  6. @berryxia70

    OpenAI 发布 Codex Python SDK,通过 pip install openai-codex 安装后,开发者可在 Python 代码中启动线程、运行 turn、实时流式输出进度、恢复会话、传图片并精细控制 sandbox 访问权限。

    引用Vaibhav (VB) Srivastav (@reach_vb)@reach_vb

    We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install openai-codex Go build with it!!

    推荐理由:SDK 把 Codex 从浏览器里的 IDE 变成可嵌进脚本的可编程基础设施,并复用现有认证。

  7. @berryxia66

    OpenAI 公布 Codex 每周活跃用户已超过 500 万,比二月份桌面 App 刚上线时增长 6 倍多。报告称知识工作者采用速度是开发者的 3 倍以上,占用户总数 20%,其中 72% 每周用它产出文档、备忘录、图像、音频或视频,增长最快的是数据分析(周环比 110%)、研究(37%)和知识产物制作(36%)。

    引用OpenAI Newsroom (@OpenAINewsroom)@OpenAINewsroom

    Codex now has more than 5M weekly active users. But the bigger story is what people are using it for: not just writing code, but getting more work done across research, analysis, content, and operations. Our new report on how Codex is becoming a productivity tool for knowledge work: openai.com/index/codex-for-k…

    推荐理由:OpenAI 公布 Codex 用户用途数据,知识工作者采用速度已超过开发者,可观察 AI 生产力工具的实际使用迁移。

  8. @vista870

    OpenAI 发布 Codex Python SDK,可用 pip install openai-codex 安装,把 Codex 直接嵌入 Python 应用和工作流。该 SDK 支持启动线程、运行轮次、流式输出进度、恢复会话、传入图片和控制沙箱权限,并可复用现有 Codex 登录态。转发的作者认为这相当于把顶级编程和生图 Agent 内置到自己的代码中,并称复用 Codex 登录态最为关键。

    引用Vaibhav (VB) Srivastav (@reach_vb)@reach_vb

    We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install openai-codex Go build with it!!

    推荐理由:原文给出 Codex Python SDK 的安装命令与复用登录态方式,读者可据此把它嵌入 Python 应用和工作流。

6月2日周二