跳到正文

多模态

文本之外的能力:视觉理解、图文混合、音视频输入输出的模型与产品进展。

当前仅显示精选新闻
236条精选相关主题图像生成AI 视频语音与音频

最新精选

第 61–80 条 · 共 236 条
8月26日周三
  1. @Alibaba_Qwen75

    通义千问发布多模态 MoE 模型 Qwen3.8-Flash 并开放权重,QwenCloud 上的 API 同步上线。模型为 125B 参数加 51B N-gram embeddings,每 token 仅激活 6B,采用 GDN + QSA 混合注意力、Gated Residual、N-gram Embedding 与 Muon 优化器,官方称其为 Qwen4 架构的早期预览,训练成本仅为 Qwen3.7-Plus 的 1/9。原生上下文 262K 可经 YaRN 扩展到 1M,评测得分 DeepSWE 1.1 58.7、SWE-bench Pro 62.5、CoWorkBench 73.9、AndroidWorld 84.5、MathVision 95.7,官方同时开放了 Qwen3.8-Flash-Next 的权重。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:训练成本降到 Qwen3.7-Plus 的 1/9 且官方称全面超越,读者可借此观察 Qwen4 新架构的取舍。

  2. @omarsar076

    阿里 Qwen 团队开放 Qwen3.8-Flash 权重,该模型为多模态 MoE,总参数 125B、每 token 激活 6B,并带 51B N-gram 嵌入,官方称其为 Qwen4 架构的早期预览。生产版本将上线 QwenCloud API,输入 $0.16/1M tokens、输出 $0.47/1M tokens,官方还给出 DeepSWE 1.1 58.7、SWE-bench Pro 62.5 等成绩。作者 @omarsar0 认为随附的技术报告比发布本身更值得读。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:发布信息列出了 Qwen3.8-Flash 的参数量、激活规模与定价,可作为判断高效多模态 MoE 路线的具体参照。

  3. @Yuchenj_UW77

    GLM-5.3-Flash(Ox Alpha)发布,320B-A18B 规模不到 GLM-5.2 的一半,却在各项基准上全面超过 GLM-5.2。该模型原生多模态、支持 1M-token 上下文窗口,以 MIT 许可发布,此前以 Ox Alpha 名义预览并完全运行在中国 AI 芯片上,权重、API、Coding Plan、ZCode、Chat、AutoClaw 等官方入口已开放。Databricks 的 Yuchen Jin 表示将尽快把 GLM-5.3-Flash 提供给客户并让它跑得很快。

    引用@Zai_org@Zai_org

    Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now across all official platforms: Weights: https://t.co/9LRMahY9Wa API: https://t.co/VcaQnzYmS9 Coding Plan: https://t.co/Nk8Y98HNhU ZCode: https://t.co/Peepqv4XSx Chat: https://t.co/WCqWT0qCQb AutoClaw: https://t.co/aGEG5HqTTb

    推荐理由:原文对比了 GLM-5.3-Flash 与 GLM-5.2 的参数量和基准成绩,可据此了解高效小模型的进展。

  4. @kimmonismus66

    Zai 发布 GLM-5.3 Flash(代号 Ox Alpha),这是一款 320B MoE 模型,每 token 仅激活 18B 参数,开放权重、MIT 许可、原生多模态、1M 上下文。据 Zai 公布的数据,它在 Terminal-Bench 2.1 得 84.3,接近 Claude Opus 4.8 的 85.0;DeepSWE 得 63.4、AutomationBench 得 48.8,均为对比中的领先成绩,并在全部六项基准上超过更大的 GLM-5.2,而服务成本仅为后者的十分之一。原文同时提醒,18B 激活参数并不等于可本地运行的 18B 模型,全部 320B 权重仍需存储。

    引用@kimmonismus@kimmonismus

    GLM-5.3 Flash ("Ox Alpha") official: Benchmarks attached. This looks exceptional for its size! GLM-5.3-Flash might be one of the most impressive efficiency releases yet. It is a 320B MoE with only 18B parameters active per token, yet Zai reports: - 84.3 on Terminal-Bench 2.1, nearly matching Claude Opus 4.8 at 85.0 - 63.4 on DeepSWE, ahead of Opus 4.8 and DeepSeek V4 Vision Exp - 48.8 on AutomationBench, ahead of Opus 4.8 and GPT-5.6 Terra - The highest GDPval-AA v2 score in its comparisonIt also beats the much larger GLM-5.2 across all six reported benchmarks while costing one-tenth as much to serve. Open weights, MIT licensed, natively multimodal, 1M context. Important caveat: 18B active parameters does not make it a normal local 18B model. All 320B weights still need to be stored. But in terms of intelligence per active parameter, this looks exceptional!

    推荐理由:原文列出六项基准对比与 MIT 许可信息,读者可据此判断这一小激活参数模型的性价比。

  5. IT Home76

    智谱开源 GLM-5.3-Flash 原生多模态模型,限时折扣价为 GLM-5.3 的 1/20

    智谱上线并开源 GLM-5.3-Flash(320B-A18B),这是 GLM-5 系列首个原生多模态模型,总参数量 320B、激活参数仅 18B。其在 Artificial Analysis Intelligence Index 取得 57 分,与 Claude Opus 4.8 持平,自研 Z.ai Code Bench 体感评估中编程表现也与之相当。

    推荐理由:320B 总参数仅激活 18B 的架构设计搭配限时 1/20 定价,可供判断开源前沿模型的成本竞争区间。

  6. Z.ai72

    智谱(Z.ai)发布 GLM-5.3-Flash,称具备有竞争力的价格与原生多模态能力,上下文窗口为 1M token,为 320B-A18B 模型并以 MIT License 开源权重。该模型此前曾以 Ox Alpha 名义预览,完全运行于中国 AI 芯片;现已在官方平台提供权重、API、Coding Plan、ZCode、Chat 和 AutoClaw 入口。

    推荐理由:官方公告同时给出价格定位、开源权重和芯片适配信息,读者可以据此评估它在现有工作流中的替换可能。

  7. @testingcatalog76

    阿里发布 Qwen3.8 Flash,一款 125B 参数的多模态 MoE 模型,原生上下文 262K,可通过 YaRN 扩展至 1M。QwenCloud API 定价为每 1M 输入 tokens 0.16 美元、每 1M 输出 tokens 0.47 美元。该模型基于新架构,是 Qwen4 所用架构的前身,在 DeepSWE 1.1 得 58.7 分、SWE-bench Pro 得 62.5 分。

    引用@Alibaba_Qwen@Alibaba_Qwen

    Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.

    推荐理由:原文给出上下文长度、API 定价与多项编码基准分数,读者可据此对比同表内 DeepSeek 与 Claude 模型的定位。

  8. 机器之心 · 微信公众号77

    阿里发布 Qwen3.8-Flash,同步开源 Qwen3.8-Flash-Next

    阿里发布 Qwen3.8-Flash,并在 Hugging Face 与 ModelScope 开放同一模型的 Qwen3.8-Flash-Next 权重,主模型 125B 参数、每 token 仅激活 6B,千问 AI 平台定价为每百万 token 输入 1 元、输出 3 元。

    推荐理由:文章拆解了 Qwen3.8-Flash 在注意力、残差与嵌入上的四处架构改动,并把它放进每任务成本的行业对比框架里。

  9. Qwen Blog69

    Qwen3.8-Flash-Next 开源,多模态 MoE 架构预览 Qwen4

    千问团队开源 Qwen3.8-Flash-Next 权重,这是一个多模态 MoE 模型,也是 Qwen4 所用架构的早期预览。文中称其角色类似 Qwen3-Next 之于 Qwen3.5,当时的混合 Gated DeltaNet + Gated Attention 设计已用于 Qwen3.5 至 Qwen3.8 系列。

    推荐理由:官方开源权重并定位为 Qwen4 架构预览,读者可据此追踪千问后续系列的架构走向。

8月25日周二
  1. @alibaba_cloud68

    Wan 3.0 已在 OpenRouter 上线,支持从文本、图像或参考素材生成 2–30 秒视频,最高 1080p。OpenRouter 公布的上线定价为 480p、720p、1080p 每秒 $0.05、$0.10、$0.20,全分辨率限时 15% 折扣。阿里云同时给出 Model Studio 与 Qwen Cloud 两个 API 接入入口。

    引用@OpenRouter@OpenRouter

    Alibaba's Wan 3.0 from @Alibaba_Wan and @alibaba_cloud is now live on OpenRouter. Generate 2-30 second videos from text, images, or references at 480p, 720p, or 1080p. Launch pricing: $0.05/$0.10/$0.20 per second, with a limited-time 15% discount across every resolution. https://t.co/yzQNUKAo3o

    推荐理由:官方给出 2–30 秒、最高 1080p 的生成规格与 OpenRouter 接入入口,可据此评估视频生成的调用成本。

  2. @alibaba_cloud69

    Wan 3.0 已在 Runway 上线,可生成视频和音频,并支持输入多张图像、视频和音频作为参考,官方称这让每次生成获得更多控制。该模型还可通过 Model Studio 和 Qwen Cloud 获取 API 访问。

    引用@runwayml@runwayml

    WAN 3.0 is now on Runway. Generate state-of-the-art video and audio with multiple image, video and audio reference inputs. Try it now at the link below. https://t.co/VzTkPZR3cz

    推荐理由:Wan 3.0 上线 Runway 并给出 API 入口,读者可了解其多模态参考输入带来的生成控制方式。

8月22日周六