跳到正文

模型发布

新模型的发布、开源与迭代:大模型厂商的旗舰更新、开源权重放出、性能与价格变化的第一时间记录。

当前仅显示精选新闻
574条精选相关主题产品更新论文研究开源生态

最新精选

第 201–220 条 · 共 574 条
8月26日周三
  1. @kimmonismus68

    Zai 发布 GLM-5.3 Flash(Ox Alpha),320B MoE 每 token 仅激活 18B 参数,采用 MIT 开源许可,原生多模态并支持 1M 上下文。

    引用@kimmonismus@kimmonismus

    The upcoming Ox Alpha is GLM-5.3 Flash (as expected): 320b total parameters, 18b active. Outperforming GLM-5.2 at 1/10th of its price and approaching Opus 4.8 on coding and agentic benchmarks. Big things incoming! https://t.co/xZZmOD1Ghu https://t.co/H2fZ9idgVa

    推荐理由:原文列出六项基准数据与 MIT 开源、1M 上下文等规格,便于读者判断这一稀疏 MoE 的效率定位。

  2. IT Home76

    智谱开源 GLM-5.3-Flash 原生多模态模型,限时折扣价为 GLM-5.3 的 1/20

    智谱上线并开源 GLM-5.3-Flash(320B-A18B),这是 GLM-5 系列首个原生多模态模型,总参数量 320B、激活参数仅 18B。其在 Artificial Analysis Intelligence Index 取得 57 分,与 Claude Opus 4.8 持平,自研 Z.ai Code Bench 体感评估中编程表现也与之相当。

    推荐理由:320B 总参数仅激活 18B 的架构设计搭配限时 1/20 定价,可供判断开源前沿模型的成本竞争区间。

  3. Z.ai72

    智谱(Z.ai)发布 GLM-5.3-Flash,称具备有竞争力的价格与原生多模态能力,上下文窗口为 1M token,为 320B-A18B 模型并以 MIT License 开源权重。该模型此前曾以 Ox Alpha 名义预览,完全运行于中国 AI 芯片;现已在官方平台提供权重、API、Coding Plan、ZCode、Chat 和 AutoClaw 入口。

    推荐理由:官方公告同时给出价格定位、开源权重和芯片适配信息,读者可以据此评估它在现有工作流中的替换可能。

  4. @AYi_AInotes76

    Qwen 团队开源 Qwen3.8-Flash,总参数 125B 加 51B N-gram 嵌入,每 token 仅激活 6B,训练成本为 Qwen3.7-Plus 的 1/9。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:Qwen3.8-Flash 以 125B 总参数、6B 激活和 1/9 训练成本给出开源 MoE 的效率样本,可对照其基准数据看架构取舍。

  5. @testingcatalog76

    阿里发布 Qwen3.8 Flash,一款 125B 参数的多模态 MoE 模型,原生上下文 262K,可通过 YaRN 扩展至 1M。QwenCloud API 定价为每 1M 输入 tokens 0.16 美元、每 1M 输出 tokens 0.47 美元。该模型基于新架构,是 Qwen4 所用架构的前身,在 DeepSWE 1.1 得 58.7 分、SWE-bench Pro 得 62.5 分。

    引用@Alibaba_Qwen@Alibaba_Qwen

    Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.

    推荐理由:原文给出上下文长度、API 定价与多项编码基准分数,读者可据此对比同表内 DeepSeek 与 Claude 模型的定位。

  6. 机器之心 · 微信公众号77

    阿里发布 Qwen3.8-Flash,同步开源 Qwen3.8-Flash-Next

    阿里发布 Qwen3.8-Flash,并在 Hugging Face 与 ModelScope 开放同一模型的 Qwen3.8-Flash-Next 权重,主模型 125B 参数、每 token 仅激活 6B,千问 AI 平台定价为每百万 token 输入 1 元、输出 3 元。

    推荐理由:文章拆解了 Qwen3.8-Flash 在注意力、残差与嵌入上的四处架构改动,并把它放进每任务成本的行业对比框架里。

  7. @kimmonismus80

    Qwen3.8-Flash-Next 发布,采用 125B MoE 参数加 51B N-gram embeddings,每 token 仅激活 6B 参数。

    引用@kimmonismus@kimmonismus

    Qwen 3.8 Flash-Next official released: A 6B-active open model just beat Claude Opus 4.6 Max across 8 of 9 comparable benchmarks! Qwen3.8-Flash-Next is a highly sparse MoE: • 125B model parameters • 51B additional n-gram embeddings • Only 6B parameters active per token It scores: • 62.5 SWE-bench Pro • 81.0 SWE-bench Multilingual • 73.9 CoworkBench • 55.7 JobBench • 73.5 Toolathlon • 81.3 IFBench • 91.7 GPQA Diamond • 91.9 LiveCodeBench It also outperforms Qwen3.8-27B and DeepSeek-V4-Flash across most of the table. Super cool release!!

    推荐理由:原文给出四项架构改动与 1/9 训练成本的对比,读者可以了解高稀疏 MoE 如何压低单 token 计算量。

  8. Qwen Blog69

    Qwen3.8-Flash-Next 开源,多模态 MoE 架构预览 Qwen4

    千问团队开源 Qwen3.8-Flash-Next 权重,这是一个多模态 MoE 模型,也是 Qwen4 所用架构的早期预览。文中称其角色类似 Qwen3-Next 之于 Qwen3.5,当时的混合 Gated DeltaNet + Gated Attention 设计已用于 Qwen3.5 至 Qwen3.8 系列。

    推荐理由:官方开源权重并定位为 Qwen4 架构预览,读者可据此追踪千问后续系列的架构走向。

8月25日周二
  1. Hugging Face Blog63

    IBM 发布 Granite 4.2 推理模型家族并详解构建过程

    IBM 发布 Granite 4.2 密集 decoder-only 推理模型家族,含 3B、8B、30B 三个规格,基于 Granite-4.1 基座(约 15T tokens 预训练,上下文窗口扩至 512K),经 SFT 与多阶段 GRPO 强化学习训练,全部以 Apache 2.0 许可开源。

    推荐理由:IBM 官方详解 Granite 4.2 训练全程,从五阶段预训练到多阶段 RL 课程,可复用的训练细节较完整。

  2. @alibaba_cloud68

    Wan 3.0 已在 OpenRouter 上线,支持从文本、图像或参考素材生成 2–30 秒视频,最高 1080p。OpenRouter 公布的上线定价为 480p、720p、1080p 每秒 $0.05、$0.10、$0.20,全分辨率限时 15% 折扣。阿里云同时给出 Model Studio 与 Qwen Cloud 两个 API 接入入口。

    引用@OpenRouter@OpenRouter

    Alibaba's Wan 3.0 from @Alibaba_Wan and @alibaba_cloud is now live on OpenRouter. Generate 2-30 second videos from text, images, or references at 480p, 720p, or 1080p. Launch pricing: $0.05/$0.10/$0.20 per second, with a limited-time 15% discount across every resolution. https://t.co/yzQNUKAo3o

    推荐理由:官方给出 2–30 秒、最高 1080p 的生成规格与 OpenRouter 接入入口,可据此评估视频生成的调用成本。

  3. @alibaba_cloud69

    Wan 3.0 已在 Runway 上线,可生成视频和音频,并支持输入多张图像、视频和音频作为参考,官方称这让每次生成获得更多控制。该模型还可通过 Model Studio 和 Qwen Cloud 获取 API 访问。

    引用@runwayml@runwayml

    WAN 3.0 is now on Runway. Generate state-of-the-art video and audio with multiple image, video and audio reference inputs. Try it now at the link below. https://t.co/VzTkPZR3cz

    推荐理由:Wan 3.0 上线 Runway 并给出 API 入口,读者可了解其多模态参考输入带来的生成控制方式。

8月24日周一
  1. @OpenRouter65

    OpenRouter 宣布隐身模型 ox-alpha 在其平台上线,单日 token 用量接近 6 万亿。配图显示该模型上线前三日 token 量达 11.6T,是 OpenRouter 历史上规模最大的模型发布,为第二名发布量的 2.6 倍。开发者可通过 ori 在编码智能体中调用,命令为 `ori [your favorite harness] --model stealth/ox-alpha`。

    推荐理由:配图给出该模型与历史发布横向对比的 token 用量,读者可据此判断它的实际调用规模。

  2. 机器之心 · 微信公众号78

    斯坦福教授 Percy Liang 团队启动 Marin 535B-A23B 模型训练并全程公开

    斯坦福大学教授 Percy Liang 与 Marin 开放实验室启动 Marin 535B-A23B 模型训练,并全程公开训练过程。该模型有 5350 亿总参数、230 亿激活参数,准备 18.75 万亿 token 训练数据,部署 11 套 GB200 NVL72 约 792 颗 GB200 GPU,预计连续训练约 3 个月,总训练计算量约 2.7e24 FLOPs。

    推荐理由:训练数据配比、实时 loss 曲线与实验日志全程公开,为观察前沿大模型训练过程提供了少见的一手材料。