推荐理由:原文列出榜单分数、参数规格与 API 价格,读者可据此判断该模型在编码场景中的性价比位置。
AI 编码
全部主题AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。
当前仅显示精选新闻最新精选
第 121–140 条 · 共 438 条@AYi_AInotes@AYi_AInotes精选AI 评分7575 
@AYi_AInotes@AYi_AInotes精选AI 评分6767
引用@Alibaba_Qwen@Alibaba_Qwen🚀Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902! 2.4T parameters. 1M context tokens. Built for real world complexity. Further post trained on Coding & Cowork, Qwen3.8-Max-0902 now delivers stronger performance across complex enterprise tasks, scientific research, and long horizon workflows. 💰Pricing per 1M tokens: $2 input, $6 output. $0.17 explicit cache hit, $0.25 implicit cache hit. Now live via API on QwenCloud. Come try it! 🙌 API: https://t.co/dq3WgMk980
推荐理由:列出新版 Qwen 在代码榜单的排位与每百万 tokens 定价,可据此评估替换现有 Agent 底座的成本。
@Alibaba_Qwen@Alibaba_Qwen精选AI 评分7272 通义千问宣布 Qwen3.8-Max-0902 在 Code Arena 总榜排名第一,并位于每百万 token 5 美元价格档的帕累托前沿。该模型可在 QwenCloud 上试用。
推荐理由:官方给出 Code Arena 登顶与每百万 token 5 美元的价格位置,可据此判断编码模型的性价比区间。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7676 
推荐理由:对比其三个月前 $26B 估值和 $492M 年化收入的融资,可以看出这轮估值与收入的同步抬升节奏。
@alibaba_cloud@alibaba_cloud精选AI 评分7373 

推荐理由:官方给出参数规模、上下文长度与分档定价,读者可据此评估它在长周期企业任务中的成本与适用性。
@Alibaba_Qwen@Alibaba_Qwen精选AI 评分6969 

推荐理由:官方同时给出参数规模、1M 上下文与 API 定价,便于开发者评估长任务场景的接入成本。
@AISafetyMemes@AISafetyMemes精选AI 评分7777
引用@claudeai@claudeaiAcross our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. https://t.co/aSb72LSxee
推荐理由:表格把 Fable 5.1 与 Fable 5 放在同一组基准上对比,可直观看到两代之间的分数差距。
IT Home精选AI 评分7676 Anthropic 发布 Claude Fable 5.1 和 Mythos 5.1:性能超越前代,缓存读取费用下调 75%
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两款模型采用相同基础模型,区别在于安全防护等级,Fable 5.1 面向所有用户开放,Mythos 5.1 仅通过可信访问计划向经审核的网络安全和生命科学机构提供。
推荐理由:两款同源模型以安全等级区分受众,基准与缓存降价的对比可供判断编程与知识工作的选型与成本。
@omarsar0@omarsar0精选AI 评分6969 
推荐理由:论文报告 openJiuwen 在固定模型策略下靠运行时可适应机制取得基准提升,读者可比对静态与动态 harness 的设计差异。
@bcherny@bcherny精选AI 评分6969 企业版、API 和 SDK 客户的价格已下调,Fable 5.1 的缓存读取从每百万 token 1 美元降至 0.25 美元。一次典型 Claude Code 会话的成本最多可降低 38%。
推荐理由:官方给出缓存读取单价和典型会话成本的变化,读者可据此估算自己的调用开销。
@testingcatalog@testingcatalog精选AI 评分7171 
引用@claudeai@claudeaiWe’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3
推荐理由:两款新模型在 Terminal-Bench 系列基准上的分数对比,可直观看出相对前代的提升幅度与推送状态。
Gemini API 更新日志精选AI 评分6262 Gemini 3.8 Flash 正式发布 GA 版本
Google 发布 gemini-3.8-flash 正式版(GA),称其为此前最智能的 Flash 模型。该模型面向长程软件工程、自主智能体和企业复杂工作流,可通过 Gemini API 模型页面和最新模型指南上手。
推荐理由:官方宣布 gemini-3.8-flash 转正为 GA 版本,定位长程软件工程与智能体场景,可据此评估是否替换现有 Flash 调用。
@rohanpaul_ai@rohanpaul_ai精选AI 评分6868
引用@finkd@finkdMuse Code is out of beta and now built to handle bigger, more complex engineering tasks. Developers can get started with one command today: curl -fsSL https://t.co/0RApZrEJMv | bash
推荐理由:原文梳理了 Muse Code 的 SDK、多子智能体并行与 rewind 回退机制,可据此判断其作为编码智能体运行时的协作与容错能力。
Anthropic Newsroom精选AI 评分8686 Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两者为同一模型、安全防护级别不同,Fable 5.1 全面开放,Mythos 5.1 限可信访问计划。
推荐理由:原文给出新旧模型基准对比、cache reads 降价幅度和访问政策变化,读者可据此评估升级成本与适用场景。
Claude Platform release notes精选AI 评分7171 Anthropic 发布 Claude Fable 5.1 与 Claude Mythos 5.1,默认 1M token 上下文并下调缓存读取价格
Anthropic 发布 Claude Fable 5.1(claude-fable-5-1),面向长时运行的智能体编码、知识工作和研究,并向 Project Glasswing 参与者提供 Claude Mythos 5.1。
推荐理由:官方发布说明列出了定价、上下文窗口和多项 API 变更,方便现有使用方评估迁移和缓存成本影响。
AI寒武纪 · 微信公众号精选AI 评分8080 Uber 公开内部 AI 软件工厂:超七成 PR 交给 Agent,每会话成本降 52%
Uber 内部超过七成的代码 PR 已由 AI Agent 完成,工程师构建了 3600 多个 Agent 技能,每天执行超过 3 万次。
推荐理由:原文拆解了 Uber 把七成 PR 交给 Agent 后仍能压低成本的具体做法,包括模型选型、上下文压缩与成本追踪体系。
@AYi_AInotes@AYi_AInotes精选AI 评分6868
引用@TencentHunyuan@TencentHunyuan🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog:https://t.co/rbl1IWRk3C HuggingFace:https://t.co/mE9wevH5XR Github:https://t.co/pyl9zckpoL https://t.co/4iW6gSuZKr
推荐理由:混元 Hy4 preview 以 770B 参数和 1M 上下文开源,Arena 代码榜排名给出编码能力的横向参照。
@AYi_AInotes@AYi_AInotes精选AI 评分6868 
推荐理由:Uber 公开了用量涨近 10 倍而账单不涨的具体工程做法,其中子 Agent 分工与工具按需挂载等杠杆可直接参考。
@testingcatalog@testingcatalog精选AI 评分6767 
引用@thsottiaux@thsottiauxWe are reseting usage for all paid users of Codex and ChatGPT Work. Please continue reading for an update on Codex usage limits. The team has been working around the clock, going through thousands of reports and shipping fixes. Depending on how you use Codex, you should see your usage go between 10% and 50% further than before. We really went with a fine comb, with many uncovered small things being longstanding and here is what we found and fixed: - Compaction. We were keeping old images during compaction, sometimes making the context large enough to trigger compaction again. After the fix, usage dropped around 10% for users making heavy use of images. Fixed. - Memory. Background memory workers could inherit Stop hooks and keep running when the hook wouldn’t let them finish. This affected fewer than 1% of users, with the long tail being pretty bad and we saw one example thread check whether it could stop 15,000 times. Fixed. - Goals. In some cases, a set /goal could finish and then keep going past the intended stop condition, or the model would keep retrying broken tools without stopping. We saw examples consume anywhere from 15% to 70% of a weekly allowance. Fixed. - Automations. Some custom schedules could run more frequently than configured. Fixed. - Subagents. Smaller models (e.g. Luna) sometimes picked more capable helpers without being explicitly asked. The same was true where the orchestrating model not running in /fast mode could request sub-agents to run /fast. Fixed. - Computer History. The older implementation could lead to repeatedly summarizing overlapping activity. For some cases we saw it consume up to one fifth of the weekly usage per week. Fixed. - Rolling task summaries. Ordinary turns were triggering extra background requests. These added about 1% to token usage. Small each time, but it adds up. We have disabled this. - MCP. Some tool results could be encoded twice. We also found tool instructions getting cut off and fetched again. Fixed. We’ve also made architectural changes to prevent these from regressing and our teams will get paged if it happens regardless. We are also working on showing you directly in the app where your usage goes so you don’t have to guess. Goes without saying that we’re resetting usage limits and I hope you enjoy a very nice Saturday!
推荐理由:原文逐项列出八类用量异常的原因与修复结果,读者可据此判断 Codex 付费额度实际能多用多少。
@thsottiaux@thsottiaux精选AI 评分7272 引用@mntruell@mntruellWe’re sorry to see that OpenAI put out a note saying they plan to block Cursor users from accessing OpenAI models in three months. OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this. Cursor was one of the very first users of OpenAI, we’ve worked closely with their team for years, and we’ve trusted their platform to be neutral infrastructure for our business.
推荐理由:OpenAI 员工反驳 Cursor 给出的 5% 流量占比,提出 token 用量不等同收入与价值,可供理解这场分歧。