跳到正文

AI 编码

AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。

当前仅显示精选新闻
438条精选相关主题Agent 智能体Cursor教程实践

最新精选

第 101–120 条 · 共 438 条
9月4日周五
  1. @steipete76

    OpenAI 公布 GPT-6 Astra,称用户在电脑上能做的事 Astra 都能代劳,且速度快。@steipete 表示已使用 Astra 数周,认为它表现好且主动,在 vitest、tsx 或 SwiftPM 上有多个 PR,Astra 在调试 OC 时发现并修复了上游依赖中的问题。

    引用@OpenAI@OpenAI

    This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast. https://t.co/gDd0IsewJw

    推荐理由:早期使用者的多周实测记录,具体列出了 Astra 在依赖调试中完成的工作,可作观察其工程能力的一手样本。

  2. The Decoder86

    OpenAI 发布 GPT-6 Astra,称其开启 AGI 时代

    OpenAI 发布其能力最强的新模型 GPT-6 Astra,总裁 Greg Brockman 称这开启了 AGI 时代。Astra 在数学、编码和网络安全基准上领先,也是 OpenAI 首个在自身安全框架下被评为 critical 的模型。测试期间,它独立发现了两个此前未知的零日漏洞。

    推荐理由:文章给出 Astra 在数学、编码和安全上的基准表现及其安全定级,读者可据此了解前沿模型的能力宣称与风险披露方式。

  3. GitHub Blog · AI & ML60

    GitHub Copilot app 新手教程:如何同时运行多个 agent 会话

    GitHub 官方博客发布面向新手的教程,介绍如何在 GitHub Copilot app 中同时运行多个 agent 会话。每个会话可运行在独立的 Git worktree 上互不干扰,各自保留上下文,用户通过 sessions 视图追踪进度,教程以 tailspin-toys 仓库演示了并行添加功能、无障碍审查和运行测试的做法。

    推荐理由:官方教程讲清了并行 agent 会话如何借助 Git worktree 隔离运行,读者可以照着示例直接上手尝试。

9月3日周四
  1. OpenAI News69

    OpenAI 发布 GPT-6 Astra 模型

    OpenAI 发布 GPT-6 Astra,称其是目前最智能、对齐程度最高的模型,在计算机操作、编码、网络安全和科学领域达到 SOTA。公告只给出能力方向,未提及发布时间、参数量或可用范围等细节。

    推荐理由:官方公告列出 GPT-6 Astra 在计算机操作、编码、网络安全与科学四类能力上的定位,读者可据此了解其能力覆盖范围。

  2. @AYi_AInotes65

    Google 发布 Gemini 3.8 Flash,官方称其在智能体和编码能力上继续提升,是 6 周内第三个更新的 Flash 模型。作者对比称其价格约为 Claude Opus 5 的 1.5 折,在法律 10.0% 对 6.7%、长视频 87.8% 对 75.4% 上反超;但 OSWorld 电脑操作 59.0% 对 75.4%、通用 Agent 规划 19.1% 对 51.8% 仍落后。

    引用@OfficialLoganK@OfficialLoganK

    Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think! https://t.co/Cj07lCBtp8

    推荐理由:文中列出 Gemini 3.8 Flash 与 Opus 5 的多项基准和价格对比,读者可据此看到廉价模型与旗舰模型当下的分工边界。

  3. Hugging Face Blog69

    Hugging Face 发布开源工具 funes,为编码 Agent 提供本地自有记忆层

    Hugging Face 发布开源工具 funes,把机器上已有的 Agent 会话记录变成可检索的记忆层,一条 funes add 命令即可接入 Claude Code、Codex、pi 和 Hermes。

    推荐理由:原文给出 funes 的本地检索管线、跨机器同步与 token 成本对比数据,读者可以据此判断它能否改善多机多 Agent 的工作流。

  4. @rohanpaul_ai67

    Meta 发布 Muse Spark 1.3,称相比 1.2 工具调用减少 20%、生成 token 减少 25%,长编码任务可用更少的模型动作和输出来完成同一交付物。该版本在 MRCR 256K–512K 上得分 98.5,超过 GPT-5.6 Sol;在 JobBench(64.9 对 65.7)、OSWorld(66.9 对 68.3)和 AutomationBench(49.4 对 50.3)上接近 Opus 5。转发的官方公告称该版本当天在 Muse Code 和 API 上线,并预告开放权重版本即将发布。

    引用@finkd@finkd

    Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API. Next up 🍉 and Muse Spark open weights releases coming soon. https://t.co/XQQEDEJGD7

    推荐理由:对比表格列出 Muse Spark 1.3 与 GPT-5.6 Sol、Opus 5 在长上下文和智能体基准上的差距,可看出其性能定位。

  5. @fofrAI74

    Google DeepMind 发布两款 Gemini 新模型:3.8 Flash 与 3.8 Flash Cyber。3.8 Flash 在软件工程、智能体任务和多步推理上较 3.7 Flash 有明显提升,3.8 Flash Cyber 则主打前沿水平的漏洞检测与自动修补。

    引用@GoogleDeepMind@GoogleDeepMind

    Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level vulnerability detection and automated patching.

    推荐理由:Google 一次发布两款 Gemini 3.8 Flash 模型,分别面向智能体任务与代码安全,可对照 3.7 Flash 看提升方向。

9月2日周三
  1. @omarsar067

    Google DeepMind 发布 Gemini 3.8 Flash 和 Gemini 3.8 Flash Cyber 两款模型。前者在软件工程、智能体任务和多步推理上较 3.7 Flash 有明显提升,后者面向网络安全,具备前沿级漏洞检测与自动修补能力。Elvis Saravia 称 Gemini 3.8 Flash 的定价有吸引力,并认为 3.8 Flash Cyber 在修补能力上处于 Pareto 前沿。

    引用@GoogleDeepMind@GoogleDeepMind

    Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level vulnerability detection and automated patching.

    推荐理由:随推附上的对比表把 Gemini 3.8 Flash 的定价与多项基准成绩与 Claude、GPT-5.6 并列,便于横向比较。

  2. @OfficialLoganK65

    Gemini 3.8 Flash 发布,主打智能体与编码能力提升,是 6 周内第三个更新的 Flash 模型。推文附带的对比表显示,其输入价格 $0.75/1M tokens、输出价格 $3.75/1M tokens,Terminal-bench 2.1 得 89.4%,LBVBench 长视频理解得 87.8%(agentic)。优惠价有效期至 2026 年 12 月 31 日,2027 年 1 月 1 日起将调整为输入 $1.50/1M tokens、输出 $7.50/1M tokens。

    推荐理由:原文列出 Gemini 3.8 Flash 的价格与多项基准对比,读者可据此判断其智能体和编码能力相较前代及竞品的位置。

  3. @frxiaobei67

    Cognition 正准备以约 470 亿美元估值再融约 10 亿美元,Bloomberg 称其年化收入已超过 9 亿美元。今年 5 月它曾以 260 亿美元估值融资 10 亿美元,当时年化收入约 4.92 亿美元;该公司是 Devin 的开发商,并已接手 Windsurf 的产品、品牌、IP 和团队。

    推荐理由:Cognition 三个月内年化收入从约 4.92 亿美元增至超 9 亿美元,估值接近翻倍,可作为观察 AI Coding Agent 商业化的样本。

  4. @kimmonismus66

    Qwen3.8-Max 升级为 Qwen3.8-Max-0902,参数量 2.4T、上下文 1M tokens,已在 QwenCloud API 上线,定价为每 1M tokens 输入 $2、输出 $6,并针对 Coding 与 Cowork 做了后训练。作者称该版本在基准上已接近 Fable 5,并注意到如此明显的提升已不再伴随版本号变化,同时认为中国实验室与美国的差距正在缩小。

    引用@Alibaba_Qwen@Alibaba_Qwen

    🚀Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902! 2.4T parameters. 1M context tokens. Built for real world complexity. Further post trained on Coding & Cowork, Qwen3.8-Max-0902 now delivers stronger performance across complex enterprise tasks, scientific research, and long horizon workflows. 💰Pricing per 1M tokens: $2 input, $6 output. $0.17 explicit cache hit, $0.25 implicit cache hit. Now live via API on QwenCloud. Come try it! 🙌 API: https://t.co/dq3WgMk980

    推荐理由:材料给出 Qwen3.8-Max-0902 的参数、上下文与定价,作者由此讨论版本迭代节奏和中美实验室差距的缩小。

  5. AI寒武纪 · 微信公众号77

    Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1 双旗舰模型,缓存读取价格降 75%

    Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1 两款底层架构相同、安全防护等级不同的模型,Fable 5.1 面向大众开放,Mythos 5.1 仅通过可信访问计划向特定机构开放。

    推荐理由:缓存读取价格下调75%并给出分子设计与金星地形重建案例,便于对比新旗舰的能力与成本结构。