跳到正文

AI 编码

AI 写代码的一切:编码助手、Vibe Coding、代码模型评测与开发工作流变革。

当前仅显示精选新闻
455条精选相关主题Agent 智能体Cursor教程实践

最新精选

第 161–180 条 · 共 455 条
8月27日周四
  1. @SiliconFlowAI66

    Zai 开源发布 GLM-5.3-Flash,并已在 SiliconFlow 上线,SiliconFlow 提供 Day-0 支持。该模型为 320B 总参数、18B 激活参数,原生多模态,采用 MIT 许可。官方称其强于 GLM-5.2 且便宜 90%,在编码与 Agent 任务上接近 Opus 4.8,成本低 95% 以上。

    推荐理由:GLM-5.3-Flash 以 320B 总参、18B 激活的开源配置上线 SiliconFlow,读者可据此对比它与前代在成本上的变化。

  2. Linear Now64

    Linear 用 1000+ 个 PR 将 React 应用从 styled-components 迁移到 StyleX

    Linear 宣布完成从 styled-components 到 StyleX 的迁移,历时逾 1000 个 PR。迁移动机包括 styled-components 进入维护模式、React 18 并发渲染下的性能回归,以及团队希望为 Agent 参与代码库建立更清晰的样式边界。

    推荐理由:Linear 团队亲历者复盘超千个 PR 的迁移过程,给出可迁移的确定性工具加 Agent 加人工判断的方法和降险步骤。

  3. 量子位 · 微信公众号81

    智谱发布并开源 GLM-5.3 Flash,为 GLM-5 系列首个原生多模态模型

    智谱发布并开源 GLM-5.3 Flash,这是 GLM-5 系列首个原生多模态模型,总参数 320B、激活参数 18B,能力超过体量 753B 的 GLM-5.2。该模型在 AA 榜单拿到 57 分,与 Claude Opus 4.8 持平,限时折扣价为后者的 1/40,并已接入 ZCode 和开放 API。模型权重已在 Hugging Face 开源,承接线上真实请求的算力来自国产芯片。

    推荐理由:实测呈现了这个 320B 模型在多模态与编码任务上的表现,并交代了低价与国产芯片部署两层背景。

8月26日周三
  1. @AYi_AInotes70

    彭博确认在 OpenRouter 匿名屠榜的 Ox Alpha 是智谱 GLM 系列的新迭代,今晚将直接开源放权重。

    原始视频预览图;未保存可播放视频URL
    引用@AYi_AInotes@AYi_AInotes

    牛来大模型刚被彭博社破案了! 在 OpenRouter 上悄悄屠榜、调用量超过 DeepSeek 一倍的神秘模型 Ox Alpha, 背后竟然是智谱,更绝的是官方确认今晚直接开源权重, 狂刷了数十万亿 token、在真实 Coding 里跑出 80% 峰值,智谱这次是真的要杀疯了, 前阵子大家都在猜这个匿名模型到底是哪家大厂的马甲, 因为它在排位赛上狂刷了数十万亿 token, 开发者用脚投票把它送上了第一, 它真正狠的地方,是把 Flash 级别的速度和成本,做出了中上游前沿模型的实战能力: 1M 超长上下文、原生支持图片和视频多模态, 而且是推理优先的架构,写代码时会先做规划再调工具 社区独立跑测试,在 10 个高难度真实 Coding 任务里拿下了 80% 的过关率, 全量 DeepSWE 跑出 64.6%, 直接咬住了 Claude Opus 4.8 和 Gemini 3.7 Flash 的身位 它不是那种打虚空跑分的绝对 SOTA, 但一个成本极低、速度极快、能看懂视频还能写复杂代码的 Flash 模型, 今晚一旦把权重放出来,无论是本地部署还是第三方 API,生态杀伤力完全不可同日而语 刚发完开源顶流 GLM-5.3,转头就把这个高频 Agent 生产力核弹开源出来, 大模型的竞争,终于从比拼谁的参数更大,彻底转向谁能让开发者用最低成本把活干完了啊

    推荐理由:匿名屠榜模型被确认为智谱 GLM 新迭代并将开源权重,读者可了解它对第三方 API 与本地部署成本的影响。

  2. @Alibaba_Qwen75

    通义千问发布多模态 MoE 模型 Qwen3.8-Flash 并开放权重,QwenCloud 上的 API 同步上线。模型为 125B 参数加 51B N-gram embeddings,每 token 仅激活 6B,采用 GDN + QSA 混合注意力、Gated Residual、N-gram Embedding 与 Muon 优化器,官方称其为 Qwen4 架构的早期预览,训练成本仅为 Qwen3.7-Plus 的 1/9。原生上下文 262K 可经 YaRN 扩展到 1M,评测得分 DeepSWE 1.1 58.7、SWE-bench Pro 62.5、CoWorkBench 73.9、AndroidWorld 84.5、MathVision 95.7,官方同时开放了 Qwen3.8-Flash-Next 的权重。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:训练成本降到 Qwen3.7-Plus 的 1/9 且官方称全面超越,读者可借此观察 Qwen4 新架构的取舍。

  3. @AYi_AInotes68

    彭博社报道称,在 OpenRouter 排行榜上登顶的匿名模型 Ox Alpha 由智谱开发,智谱将开源其权重,该模型目前仍免费使用。按帖中说法,它是以推理优先架构设计的编码与智能体模型,原生支持文本、图像和视频输入,社区测试在 10 个高难度真实 Coding 任务中取得 80% 过关率,全量 DeepSWE 为 64.6%。作者还提到智谱此前刚发布开源模型 GLM-5.3。

    推荐理由:彭博社确认匿名模型 Ox Alpha 出自智谱并将在今晚开源权重,读者可了解它的真实来源与开放安排。

  4. @kimmonismus68

    Zai 发布 GLM-5.3 Flash(Ox Alpha),320B MoE 每 token 仅激活 18B 参数,采用 MIT 开源许可,原生多模态并支持 1M 上下文。

    引用@kimmonismus@kimmonismus

    The upcoming Ox Alpha is GLM-5.3 Flash (as expected): 320b total parameters, 18b active. Outperforming GLM-5.2 at 1/10th of its price and approaching Opus 4.8 on coding and agentic benchmarks. Big things incoming! https://t.co/xZZmOD1Ghu https://t.co/H2fZ9idgVa

    推荐理由:原文列出六项基准数据与 MIT 开源、1M 上下文等规格,便于读者判断这一稀疏 MoE 的效率定位。

  5. IT Home76

    智谱开源 GLM-5.3-Flash 原生多模态模型,限时折扣价为 GLM-5.3 的 1/20

    智谱上线并开源 GLM-5.3-Flash(320B-A18B),这是 GLM-5 系列首个原生多模态模型,总参数量 320B、激活参数仅 18B。其在 Artificial Analysis Intelligence Index 取得 57 分,与 Claude Opus 4.8 持平,自研 Z.ai Code Bench 体感评估中编程表现也与之相当。

    推荐理由:320B 总参数仅激活 18B 的架构设计搭配限时 1/20 定价,可供判断开源前沿模型的成本竞争区间。

  6. @AYi_AInotes76

    Qwen 团队开源 Qwen3.8-Flash,总参数 125B 加 51B N-gram 嵌入,每 token 仅激活 6B,训练成本为 Qwen3.7-Plus 的 1/9。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:Qwen3.8-Flash 以 125B 总参数、6B 激活和 1/9 训练成本给出开源 MoE 的效率样本,可对照其基准数据看架构取舍。

  7. @testingcatalog76

    阿里发布 Qwen3.8 Flash,一款 125B 参数的多模态 MoE 模型,原生上下文 262K,可通过 YaRN 扩展至 1M。QwenCloud API 定价为每 1M 输入 tokens 0.16 美元、每 1M 输出 tokens 0.47 美元。该模型基于新架构,是 Qwen4 所用架构的前身,在 DeepSWE 1.1 得 58.7 分、SWE-bench Pro 得 62.5 分。

    引用@Alibaba_Qwen@Alibaba_Qwen

    Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.

    推荐理由:原文给出上下文长度、API 定价与多项编码基准分数,读者可据此对比同表内 DeepSeek 与 Claude 模型的定位。

  8. Linear Now69

    Linear 完成 9900 万美元回购,估值翻倍至 25 亿美元,ARR 突破 1 亿美元

    Linear 完成 9900 万美元回购,估值 25 亿美元,是去年 12.5 亿美元的两倍,Accel、01A、Salesforce Ventures 和 S32 参与其中;公司现金流为正,账上现金超过历史融资总额。

    推荐理由:官方披露 2.5B 估值回购与 100M ARR 等关键数据,读者可借此了解 Linear 的经营现状与智能体业务进展。

  9. @kimmonismus71

    OpenAI 在关于 Jalapeño 芯片的博客中披露,GPT-Astra 用 Codex 编写并优化底层 kernel,两个月内把三个原本不在计划内的开源权重模型带到该芯片上的高性能。在选定的 attention 和 MoE 模块上,AI 生成的实现比现有人工专家代码快 1.5–1.8 倍,这些数字只适用于选定模块而非完整模型。引用的推文还提到,Jalapeño 在 GPT-OSS 120B、DeepSeek R1 670B 和 Kimi K2.5 1T 上实现 1.5–1.9 倍每瓦 AI 吞吐与 1.7–3.6 倍更低的端到端延迟,OpenAI 计划 2026 年底开始部署。

    引用@kimmonismus@kimmonismus

    Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI’s own InferenceX testing! For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance. The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads. OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape. Probably thats why Tibo said that in 1-2 years 750token/s will be the default

    推荐理由:原文给出 GPT-Astra 用 Codex 生成 kernel 的具体加速数据,读者可据此了解 AI 编写底层算子的表现。

8月24日周一
  1. @OpenRouter65

    OpenRouter 宣布隐身模型 ox-alpha 在其平台上线,单日 token 用量接近 6 万亿。配图显示该模型上线前三日 token 量达 11.6T,是 OpenRouter 历史上规模最大的模型发布,为第二名发布量的 2.6 倍。开发者可通过 ori 在编码智能体中调用,命令为 `ori [your favorite harness] --model stealth/ox-alpha`。

    推荐理由:配图给出该模型与历史发布横向对比的 token 用量,读者可据此判断它的实际调用规模。

  2. @alibaba_cloud74

    Kimi-K3 已在阿里云 Model Studio 和 Qwen Cloud 上线,拥有 2.8T 参数和原生 1M token 上下文,面向仓库级软件工程、视觉创作以及能自主规划下一步的智能体。阿里云同步开放 1M token 免费试用,提供 1M-token free trial 与上线即有的 autoTPM 弹性容量。

    推荐理由:阿里云给出 Kimi-K3 的参数量、上下文规格与两个平台的试用入口,可据此判断其可用渠道与定位。

8月22日周六
  1. @AYi_AInotes68

    Google 发布 Gemini 3.7 Flash,官方将其定位为迄今最智能的工作马模型,Google 高管发推称这是其史上增长最快的一次模型发布。

    引用@OfficialLoganK@OfficialLoganK

    Gemini 3.7 is our fastest growing model launch to date, amazing to see the reception!!!

    推荐理由:文章给出 Gemini 3.7 Flash 的代码评测与价格变化,并讨论低价模型对 Agent 算力成本结构的影响。