跳到正文

模型发布

新模型的发布、开源与迭代:大模型厂商的旗舰更新、开源权重放出、性能与价格变化的第一时间记录。

当前仅显示精选新闻
574条精选相关主题产品更新论文研究开源生态

最新精选

第 181–200 条 · 共 574 条
8月28日周五
  1. @testingcatalog71

    腾讯发布开放权重模型 Hy4 Preview,采用 Apache License 2.0 许可。该模型为新一代 Mixture-of-Experts(MoE)旗舰模型,770B 参数、49B 激活、1M 上下文窗口。腾讯官方博客、HuggingFace 与 GitHub 均已放出模型入口,作者表示将开始测试。

    引用@TencentHunyuan@TencentHunyuan

    🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog:https://t.co/rbl1IWRk3C HuggingFace:https://t.co/mE9wevH5XR Github:https://t.co/pyl9zckpoL https://t.co/4iW6gSuZKr

    推荐理由:腾讯以 Apache License 2.0 开放 770B 参数 MoE 模型,可对照其与现有开源模型的参数与上下文规模。

8月27日周四
  1. @SiliconFlowAI66

    Zai 开源发布 GLM-5.3-Flash,并已在 SiliconFlow 上线,SiliconFlow 提供 Day-0 支持。该模型为 320B 总参数、18B 激活参数,原生多模态,采用 MIT 许可。官方称其强于 GLM-5.2 且便宜 90%,在编码与 Agent 任务上接近 Opus 4.8,成本低 95% 以上。

    推荐理由:GLM-5.3-Flash 以 320B 总参、18B 激活的开源配置上线 SiliconFlow,读者可据此对比它与前代在成本上的变化。

  2. 量子位 · 微信公众号77

    阿里开源 Qwen3.8-Flash-Next 权重,125B MoE 外加 51B N-gram 嵌入,4090 可跑满血版

    阿里开源 Qwen3.8-Flash-Next 全部权重,这是 Qwen4 新架构的早期预览版,采用 125B 参数 MoE 配 6B 激活参数,并附加 51B 的 N-gram Embedding 参数。

    推荐理由:原文给出了新架构的权重配置与 N-gram 嵌入设计细节,可据此了解消费级硬件部署百亿级 MoE 的路径。

  3. 数字生命卡兹克 · 微信公众号77

    智谱发布 GLM-5.3-Flash,320B MoE 原生多模态模型

    智谱发布 GLM-5.3-Flash,即一周来在 OpenRouter 和 OpenCode 上匿名测试的 OX Alpha。该模型总参数 320B、每次激活 18B,是 GLM-5 系列首个原生多模态模型,支持图像、视频和文本输入,原生上下文 100 万,使用 30 万亿 Token 多模态数据预训练。

    推荐理由:模型定价与国产卡推理规模是这篇实测里最值得看的两点,读者可据此判断日常任务的模型选择。

  4. @rohanpaul_ai69

    Google 发布 Gemini 3.5 Transcribe,一款不再逐字记录、而是输出用户本意的语音转写模型,现已通过 API 开放。

    原始视频预览图;未保存可播放视频URL原始视频预览图;未保存可播放视频URL原始视频预览图;未保存可播放视频URL
    引用@sundarpichai@sundarpichai

    Say hello to Gemini 3.5 Transcribe! - Build apps that understand user speech / intent, even w/ multiple speakers! - Auto-detection of 85+ languages out of the box - Custom vocab adaptation for specialized jargon... SGTM:) API available now in @GoogleAIStudio and Gemini Enterprise, or try it in the Gemini app on macOS or Rambler on Android! More details: https://t.co/AduutCb3M7

    推荐理由:原文列出两个 API 端点各自的能力与时长限制,读者可据此判断语音转写在工作流中的可用边界。

  5. @omarsar070

    @Zai_org 发布 GLM-5.3-Flash,一个 320B-A18B 的原生多模态模型,支持 1M token 上下文窗口,采用 MIT 许可,权重、API 与编码方案等入口同步开放。该模型此前以 Ox Alpha 为名预览,完全运行在中国 AI 芯片上。Elvis Saravia 称自己在 /eli5 视觉讲解场景中使用它效果好,并推荐在 Pi 或 Hermes Agent 的新 agent playground 中试用。

    原始视频预览图;未保存可播放视频URL
    引用@Zai_org@Zai_org

    Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now across all official platforms: Weights: https://t.co/9LRMahY9Wa API: https://t.co/VcaQnzYmS9 Coding Plan: https://t.co/Nk8Y98HNhU ZCode: https://t.co/Peepqv4XSx Chat: https://t.co/WCqWT0qCQb AutoClaw: https://t.co/aGEG5HqTTb

    推荐理由:引用内容公布了 GLM-5.3-Flash 的参数量、上下文窗口与开源许可,读者可据此判断其定位与获取方式。

  6. 量子位 · 微信公众号81

    智谱发布并开源 GLM-5.3 Flash,为 GLM-5 系列首个原生多模态模型

    智谱发布并开源 GLM-5.3 Flash,这是 GLM-5 系列首个原生多模态模型,总参数 320B、激活参数 18B,能力超过体量 753B 的 GLM-5.2。该模型在 AA 榜单拿到 57 分,与 Claude Opus 4.8 持平,限时折扣价为后者的 1/40,并已接入 ZCode 和开放 API。模型权重已在 Hugging Face 开源,承接线上真实请求的算力来自国产芯片。

    推荐理由:实测呈现了这个 320B 模型在多模态与编码任务上的表现,并交代了低价与国产芯片部署两层背景。

  7. Gemini API 更新日志64

    Gemini Omni Flash 转正为 gemini-omni-1.1-flash,新增视频延伸与首尾帧插值

    Google 发布 gemini-omni-1.1-flash,即对话式视频生成与编辑模型 Gemini Omni Flash 的 GA 版本。新增视频延伸、首尾帧插值(最多 2 张图)以及 360p 到 4K 的分辨率参数,1080p 和 4K 通过超分辨率生成;预览端点 gemini-omni-flash-preview 将于 2026 年 9 月 30 日弃用。

    推荐理由:官方更新日志给出 GA 版本的新能力清单和预览端点停用时间,开发者可据此规划迁移。

8月26日周三
  1. @AYi_AInotes70

    彭博确认在 OpenRouter 匿名屠榜的 Ox Alpha 是智谱 GLM 系列的新迭代,今晚将直接开源放权重。

    原始视频预览图;未保存可播放视频URL
    引用@AYi_AInotes@AYi_AInotes

    牛来大模型刚被彭博社破案了! 在 OpenRouter 上悄悄屠榜、调用量超过 DeepSeek 一倍的神秘模型 Ox Alpha, 背后竟然是智谱,更绝的是官方确认今晚直接开源权重, 狂刷了数十万亿 token、在真实 Coding 里跑出 80% 峰值,智谱这次是真的要杀疯了, 前阵子大家都在猜这个匿名模型到底是哪家大厂的马甲, 因为它在排位赛上狂刷了数十万亿 token, 开发者用脚投票把它送上了第一, 它真正狠的地方,是把 Flash 级别的速度和成本,做出了中上游前沿模型的实战能力: 1M 超长上下文、原生支持图片和视频多模态, 而且是推理优先的架构,写代码时会先做规划再调工具 社区独立跑测试,在 10 个高难度真实 Coding 任务里拿下了 80% 的过关率, 全量 DeepSWE 跑出 64.6%, 直接咬住了 Claude Opus 4.8 和 Gemini 3.7 Flash 的身位 它不是那种打虚空跑分的绝对 SOTA, 但一个成本极低、速度极快、能看懂视频还能写复杂代码的 Flash 模型, 今晚一旦把权重放出来,无论是本地部署还是第三方 API,生态杀伤力完全不可同日而语 刚发完开源顶流 GLM-5.3,转头就把这个高频 Agent 生产力核弹开源出来, 大模型的竞争,终于从比拼谁的参数更大,彻底转向谁能让开发者用最低成本把活干完了啊

    推荐理由:匿名屠榜模型被确认为智谱 GLM 新迭代并将开源权重,读者可了解它对第三方 API 与本地部署成本的影响。

  2. @Alibaba_Qwen75

    通义千问发布多模态 MoE 模型 Qwen3.8-Flash 并开放权重,QwenCloud 上的 API 同步上线。模型为 125B 参数加 51B N-gram embeddings,每 token 仅激活 6B,采用 GDN + QSA 混合注意力、Gated Residual、N-gram Embedding 与 Muon 优化器,官方称其为 Qwen4 架构的早期预览,训练成本仅为 Qwen3.7-Plus 的 1/9。原生上下文 262K 可经 YaRN 扩展到 1M,评测得分 DeepSWE 1.1 58.7、SWE-bench Pro 62.5、CoWorkBench 73.9、AndroidWorld 84.5、MathVision 95.7,官方同时开放了 Qwen3.8-Flash-Next 的权重。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:训练成本降到 Qwen3.7-Plus 的 1/9 且官方称全面超越,读者可借此观察 Qwen4 新架构的取舍。

  3. @rohanpaul_ai66

    Z.ai 披露此前匿名上线的模型 Ox Alpha 实为 GLM-5.3-Flash,作者转述的报道称该服务由数万个国产加速器支撑,而非 NVIDIA GPU 集群。

    引用@rohanpaul_ai@rohanpaul_ai

    OxAlpha is a new iteration of GLM, from China’s Z .ai And it will change how you run long running agent fast. --- bloomberg .com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek https://t.co/ZOTlnBgOsy

    推荐理由:原文确认 Ox Alpha 即 GLM-5.3-Flash,并给出参数规模、架构改动与基准对比,便于对照其相对 GLM-5.2 的位置。

  4. @omarsar076

    阿里 Qwen 团队开放 Qwen3.8-Flash 权重,该模型为多模态 MoE,总参数 125B、每 token 激活 6B,并带 51B N-gram 嵌入,官方称其为 Qwen4 架构的早期预览。生产版本将上线 QwenCloud API,输入 $0.16/1M tokens、输出 $0.47/1M tokens,官方还给出 DeepSWE 1.1 58.7、SWE-bench Pro 62.5 等成绩。作者 @omarsar0 认为随附的技术报告比发布本身更值得读。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:发布信息列出了 Qwen3.8-Flash 的参数量、激活规模与定价,可作为判断高效多模态 MoE 路线的具体参照。

  5. @Yuchenj_UW77

    GLM-5.3-Flash(Ox Alpha)发布,320B-A18B 规模不到 GLM-5.2 的一半,却在各项基准上全面超过 GLM-5.2。该模型原生多模态、支持 1M-token 上下文窗口,以 MIT 许可发布,此前以 Ox Alpha 名义预览并完全运行在中国 AI 芯片上,权重、API、Coding Plan、ZCode、Chat、AutoClaw 等官方入口已开放。Databricks 的 Yuchen Jin 表示将尽快把 GLM-5.3-Flash 提供给客户并让它跑得很快。

    引用@Zai_org@Zai_org

    Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now across all official platforms: Weights: https://t.co/9LRMahY9Wa API: https://t.co/VcaQnzYmS9 Coding Plan: https://t.co/Nk8Y98HNhU ZCode: https://t.co/Peepqv4XSx Chat: https://t.co/WCqWT0qCQb AutoClaw: https://t.co/aGEG5HqTTb

    推荐理由:原文对比了 GLM-5.3-Flash 与 GLM-5.2 的参数量和基准成绩,可据此了解高效小模型的进展。

  6. @AYi_AInotes68

    彭博社报道称,在 OpenRouter 排行榜上登顶的匿名模型 Ox Alpha 由智谱开发,智谱将开源其权重,该模型目前仍免费使用。按帖中说法,它是以推理优先架构设计的编码与智能体模型,原生支持文本、图像和视频输入,社区测试在 10 个高难度真实 Coding 任务中取得 80% 过关率,全量 DeepSWE 为 64.6%。作者还提到智谱此前刚发布开源模型 GLM-5.3。

    推荐理由:彭博社确认匿名模型 Ox Alpha 出自智谱并将在今晚开源权重,读者可了解它的真实来源与开放安排。

  7. @kimmonismus66

    Zai 发布 GLM-5.3 Flash(代号 Ox Alpha),这是一款 320B MoE 模型,每 token 仅激活 18B 参数,开放权重、MIT 许可、原生多模态、1M 上下文。据 Zai 公布的数据,它在 Terminal-Bench 2.1 得 84.3,接近 Claude Opus 4.8 的 85.0;DeepSWE 得 63.4、AutomationBench 得 48.8,均为对比中的领先成绩,并在全部六项基准上超过更大的 GLM-5.2,而服务成本仅为后者的十分之一。原文同时提醒,18B 激活参数并不等于可本地运行的 18B 模型,全部 320B 权重仍需存储。

    引用@kimmonismus@kimmonismus

    GLM-5.3 Flash ("Ox Alpha") official: Benchmarks attached. This looks exceptional for its size! GLM-5.3-Flash might be one of the most impressive efficiency releases yet. It is a 320B MoE with only 18B parameters active per token, yet Zai reports: - 84.3 on Terminal-Bench 2.1, nearly matching Claude Opus 4.8 at 85.0 - 63.4 on DeepSWE, ahead of Opus 4.8 and DeepSeek V4 Vision Exp - 48.8 on AutomationBench, ahead of Opus 4.8 and GPT-5.6 Terra - The highest GDPval-AA v2 score in its comparisonIt also beats the much larger GLM-5.2 across all six reported benchmarks while costing one-tenth as much to serve. Open weights, MIT licensed, natively multimodal, 1M context. Important caveat: 18B active parameters does not make it a normal local 18B model. All 320B weights still need to be stored. But in terms of intelligence per active parameter, this looks exceptional!

    推荐理由:原文列出六项基准对比与 MIT 许可信息,读者可据此判断这一小激活参数模型的性价比。