跳到正文

AI 编码

AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。

当前仅显示精选新闻
438条精选相关主题Agent 智能体Cursor教程实践

最新精选

第 121–140 条 · 共 438 条
9月2日周三
  1. @AYi_AInotes67

    阿里更新旗舰模型 Qwen3.8-Max-0902,在 CodeArena WebDev 榜单以 1691 分排名第一,超过 Claude Opus 5 的 1688 分和 Kimi K3 的 1674 分,较前代 Qwen3.8-Max 的 1669 分高出 22 分。

    引用@Alibaba_Qwen@Alibaba_Qwen

    🚀Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902! 2.4T parameters. 1M context tokens. Built for real world complexity. Further post trained on Coding & Cowork, Qwen3.8-Max-0902 now delivers stronger performance across complex enterprise tasks, scientific research, and long horizon workflows. 💰Pricing per 1M tokens: $2 input, $6 output. $0.17 explicit cache hit, $0.25 implicit cache hit. Now live via API on QwenCloud. Come try it! 🙌 API: https://t.co/dq3WgMk980

    推荐理由:列出新版 Qwen 在代码榜单的排位与每百万 tokens 定价,可据此评估替换现有 Agent 底座的成本。

  2. @alibaba_cloud73

    阿里云将 Qwen3.8-Max 升级为 Qwen3.8-Max-0902,参数规模 2.4T,支持 1M token 上下文,并称其在 Coding 与 Cowork 上做了进一步后训练,在复杂企业任务、科学研究和长周期工作流上表现更强。定价为每 1M token 输入 2 美元、输出 6 美元,显式缓存命中 0.17 美元,隐式缓存命中 0.25 美元。模型可通过阿里云 Model Studio 与 Qwen Cloud 调用。

    推荐理由:官方给出参数规模、上下文长度与分档定价,读者可据此评估它在长周期企业任务中的成本与适用性。

  3. @Alibaba_Qwen69

    通义千问将 Qwen3.8-Max 升级为 Qwen3.8-Max-0902,参数规模 2.4T、上下文窗口 1M,已在 QwenCloud 上线 API。该版本在 Coding 与 Cowork 方向进一步后训练,官方称在复杂企业任务、科学研究和长周期工作流上表现更强。API 定价为每 100 万 token 输入 $2、输出 $6,显式缓存命中 $0.17、隐式缓存命中 $0.25。

    推荐理由:官方同时给出参数规模、1M 上下文与 API 定价,便于开发者评估长任务场景的接入成本。

  4. @AISafetyMemes77

    Fable 5.1 在 Terminal-Bench-Science 0.1 上得分 52.6%,超过 Fable 5 的 24.7%;在 Terminal-Bench 4.0 上以 55.8% 对 42.0% 领先,GDPval-AA v2 上为 1853 对 1723。引用内容称该模型在多项基准上刷新标准。作者指出这两代之间只相隔 2.5 个月,而不是数年。

    引用@claudeai@claudeai

    Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. https://t.co/aSb72LSxee

    推荐理由:表格把 Fable 5.1 与 Fable 5 放在同一组基准上对比,可直观看到两代之间的分数差距。

  5. IT Home76

    Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1:性能超越前代,缓存读取费用下调 75%

    Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两款模型采用相同基础模型,区别在于安全防护等级,Fable 5.1 面向所有用户开放,Mythos 5.1 仅通过可信访问计划向经审核的网络安全和生命科学机构提供。

    推荐理由:两款同源模型以安全等级区分受众,基准与缓存降价的对比可供判断编程与知识工作的选型与成本。

  6. @testingcatalog71

    Anthropic 宣布推出 Claude Fable 5.1 和 Claude Mythos 5.1,并称其面向编码与知识工作。据 Testing Catalog 引述,Fable 5.1 在 Terminal-Bench-Science 0.1 上得分 52.6%,是 Fable 5 的两倍多;在 Terminal-Bench 4.0 上得分 55.8%,Fable 5 为 42.0%。两款模型目前正在 Claude 上推送。

    引用@claudeai@claudeai

    We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3

    推荐理由:两款新模型在 Terminal-Bench 系列基准上的分数对比,可直观看出相对前代的提升幅度与推送状态。

  7. Gemini API 更新日志62

    Gemini 3.8 Flash 正式发布 GA 版本

    Google 发布 gemini-3.8-flash 正式版(GA),称其为此前最智能的 Flash 模型。该模型面向长程软件工程、自主智能体和企业复杂工作流,可通过 Gemini API 模型页面和最新模型指南上手。

    推荐理由:官方宣布 gemini-3.8-flash 转正为 GA 版本,定位长程软件工程与智能体场景,可据此评估是否替换现有 Flash 调用。

9月1日周二
  1. @rohanpaul_ai68

    Meta 的 Muse Code 结束 beta,并开放可编程能力,新 SDK 让开发者可在 Muse 的会话、工具、权限和持久状态之上构建自己的 agent 与应用。工作流可并行协调多个子智能体,跨会话消息让多个相互独立的编码任务彼此协调,而非各自孤立。rewind 功能可回退到会话中较早的安全点,同时撤销该点之后的代码改动和对话状态。

    原始视频预览图;未保存可播放视频URL
    引用@finkd@finkd

    Muse Code is out of beta and now built to handle bigger, more complex engineering tasks. Developers can get started with one command today: curl -fsSL https://t.co/0RApZrEJMv | bash

    推荐理由:原文梳理了 Muse Code 的 SDK、多子智能体并行与 rewind 回退机制,可据此判断其作为编码智能体运行时的协作与容错能力。

  2. Claude Platform release notes71

    Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1,默认 1M token 上下文并下调缓存读取价格

    Anthropic 发布 Claude Fable 5.1(claude-fable-5-1),面向长时运行的智能体编码、知识工作和研究,并向 Project Glasswing 参与者提供 Claude Mythos 5.1。

    推荐理由:官方发布说明列出了定价、上下文窗口和多项 API 变更,方便现有使用方评估迁移和缓存成本影响。

8月31日周一
8月30日周日
  1. @AYi_AInotes68

    腾讯混元发布 Hy4 preview,770B 总参数、49B 激活参数、1M 上下文,面向生产力场景并开源,官方称保持可负担的一致定价。官方同时给出博客、HuggingFace 和 GitHub 入口,并邀请用户反馈问题。转发该消息的博主称 Hy4 已进入 Arena 代码榜全球第 5。

    原始视频预览图;未保存可播放视频URL
    引用@TencentHunyuan@TencentHunyuan

    🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog:https://t.co/rbl1IWRk3C HuggingFace:https://t.co/mE9wevH5XR Github:https://t.co/pyl9zckpoL https://t.co/4iW6gSuZKr

    推荐理由:混元 Hy4 preview 以 770B 参数和 1M 上下文开源,Arena 代码榜排名给出编码能力的横向参照。

  2. @testingcatalog67

    OpenAI 再次为 Codex 和 ChatGPT Work 的全部付费用户重置用量额度,并修复多项用量消耗问题,用户整体可用量比此前多出 10% 到 50%。官方公布的修复包括 compaction 阶段遗留旧图片导致重复压缩、后台 memory worker 继承 Stop hooks 无法停止、set /goal 超出停止条件继续运行、自定义 automations 执行频率高于配置、子智能体未经要求调用更强模型,以及 Computer History 重复汇总重叠活动、普通回合触发额外后台请求、MCP 工具结果被重复编码等,并称已做架构调整防止回退。官方还表示正在开发在应用内直接展示用量去向的功能。

    引用@thsottiaux@thsottiaux

    We are reseting usage for all paid users of Codex and ChatGPT Work. Please continue reading for an update on Codex usage limits. The team has been working around the clock, going through thousands of reports and shipping fixes. Depending on how you use Codex, you should see your usage go between 10% and 50% further than before. We really went with a fine comb, with many uncovered small things being longstanding and here is what we found and fixed: - Compaction. We were keeping old images during compaction, sometimes making the context large enough to trigger compaction again. After the fix, usage dropped around 10% for users making heavy use of images. Fixed. - Memory. Background memory workers could inherit Stop hooks and keep running when the hook wouldn’t let them finish. This affected fewer than 1% of users, with the long tail being pretty bad and we saw one example thread check whether it could stop 15,000 times. Fixed. - Goals. In some cases, a set /goal could finish and then keep going past the intended stop condition, or the model would keep retrying broken tools without stopping. We saw examples consume anywhere from 15% to 70% of a weekly allowance. Fixed. - Automations. Some custom schedules could run more frequently than configured. Fixed. - Subagents. Smaller models (e.g. Luna) sometimes picked more capable helpers without being explicitly asked. The same was true where the orchestrating model not running in /fast mode could request sub-agents to run /fast. Fixed. - Computer History. The older implementation could lead to repeatedly summarizing overlapping activity. For some cases we saw it consume up to one fifth of the weekly usage per week. Fixed. - Rolling task summaries. Ordinary turns were triggering extra background requests. These added about 1% to token usage. Small each time, but it adds up. We have disabled this. - MCP. Some tool results could be encoded twice. We also found tool instructions getting cut off and fetched again. Fixed. We’ve also made architectural changes to prevent these from regressing and our teams will get paged if it happens regardless. We are also working on showing you directly in the app where your usage goes so you don’t have to guess. Goes without saying that we’re resetting usage limits and I hope you enjoy a very nice Saturday!

    推荐理由:原文逐项列出八类用量异常的原因与修复结果,读者可据此判断 Codex 付费额度实际能多用多少。

  3. @thsottiaux72

    Cursor CEO Michael Truell 表示,OpenAI 计划在三个月内阻止 Cursor 用户访问 OpenAI 模型,OpenAI 模型服务约 5% 的 Cursor 用户流量,Cursor 正与 OpenAI 团队沟通解决。OpenAI 的 Tibo 转发回应称,这个 5% 应带上强烈前提,token 既不代表收入也不代表创造的价值,较小或较弱的模型完成同一任务需要更多 token,会显著抬高流量占比。

    引用@mntruell@mntruell

    We’re sorry to see that OpenAI put out a note saying they plan to block Cursor users from accessing OpenAI models in three months. OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this. Cursor was one of the very first users of OpenAI, we’ve worked closely with their team for years, and we’ve trusted their platform to be neutral infrastructure for our business.

    推荐理由:OpenAI 员工反驳 Cursor 给出的 5% 流量占比,提出 token 用量不等同收入与价值,可供理解这场分歧。