跳到正文

千问 Qwen

阿里千问 Qwen 系列的开源发布与迭代:从旗舰模型到端侧小模型的全谱系动态。

当前仅显示精选新闻

最新精选

第 21–40 条 · 共 71 条
8月28日周五
8月27日周四
  1. @omarsar069

    一项针对自我演化编码智能体的新研究提出 EvoMal 攻击,共享技能库中的恶意技能不会被直接调用,而是被智能体当作编写模板复制、保存并执行,在库中自行扩散。在 153 个与工具相关的 SWE-bench Verified 任务、六个模型上,智能体自我投毒率为 20.3% 至 41.8%,被投毒的库最终持有的恶意技能是植入数量的 4.9 至 9.0 倍。

    推荐理由:论文量化了共享技能库中恶意技能被智能体自行复制扩散的机制,并给出一种提示词缓解办法,可供搭建技能库的团队参考。

  2. 量子位 · 微信公众号77

    阿里开源 Qwen3.8-Flash-Next 权重,125B MoE 外加 51B N-gram 嵌入,4090 可跑满血版

    阿里开源 Qwen3.8-Flash-Next 全部权重,这是 Qwen4 新架构的早期预览版,采用 125B 参数 MoE 配 6B 激活参数,并附加 51B 的 N-gram Embedding 参数。

    推荐理由:原文给出了新架构的权重配置与 N-gram 嵌入设计细节,可据此了解消费级硬件部署百亿级 MoE 的路径。

  3. @AYi_AInotes67

    作者认为「牛来」有成为超级 IP 的潜力,如果导演允许二创,社区可能搓出「牛来宇宙」,并举例超牛、豹拉大战牛来。其引用的帖子还提到,Qwen 团队这次总参数 125B、每 token 仅激活 6B,训练成本降到上一代的 1/9。

    原始视频预览图;未保存可播放视频URL
    引用@AYi_AInotes@AYi_AInotes

    Damn,125B 总参数,每 token 只激活 6B,训练成本砍到上一代的 1/9,SWE-bench Pro 62.5 超过 Claude Opus 4.6 Max 的 53.4。 Qwen 团队今天这波操作真是让我目瞪口呆,, 这根本不是什么 3.8 小更新啊, 简直就是把 Qwen4 的核心架构提前开源出来掀桌了啊, 整篇报告看完,数据简直炸裂: 总参数 125B 加上 51B 嵌入,但每个 token 实际只激活区区 6B 参数, 训练成本直接砍到上一代的九分之一, 但在 SWE-bench Pro 的编程基准里拿下 62.5,正面反超了 Claude 的旗舰模型 通义这次可以说是把大模型的架构刀法彻底改写了: 他们用类似大脑记忆的方式,平时压缩记忆,需要时才做精准检索, 1M 超长上下文的吞吐直接暴涨了 8 倍多 最绝的是加了 51B 的常用词字典,直接放在内存里查表, 几乎不占 GPU 算力,就白白吃下了巨大的模型容量 这根本不是常规的堆算力,而是智能密度和系统效率的全面革命, 故意在全量 Qwen4 发布前把架构权重开源, 就是让开源社区、vLLM 和量化生态先把底座铺满, 再加上输入 1 块钱输出 3 块钱每百万 token 的定价值,完全是要把 Agent 应用的门槛直接打穿le 以前大家总觉得模型越大越好, 现在通义用一台极速跑车告诉你,架构做对了,用极低的成本一样能跑出顶级的智能 属于架构效率狂欢的时代,真被通义给踢开大门了啊

    推荐理由:作者从角色形象、二创空间和导演态度出发推测牛来的 IP 潜力,二创若放开可能催生系列化的衍生内容。

8月26日周三
  1. @Alibaba_Qwen75

    通义千问发布多模态 MoE 模型 Qwen3.8-Flash 并开放权重,QwenCloud 上的 API 同步上线。模型为 125B 参数加 51B N-gram embeddings,每 token 仅激活 6B,采用 GDN + QSA 混合注意力、Gated Residual、N-gram Embedding 与 Muon 优化器,官方称其为 Qwen4 架构的早期预览,训练成本仅为 Qwen3.7-Plus 的 1/9。原生上下文 262K 可经 YaRN 扩展到 1M,评测得分 DeepSWE 1.1 58.7、SWE-bench Pro 62.5、CoWorkBench 73.9、AndroidWorld 84.5、MathVision 95.7,官方同时开放了 Qwen3.8-Flash-Next 的权重。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:训练成本降到 Qwen3.7-Plus 的 1/9 且官方称全面超越,读者可借此观察 Qwen4 新架构的取舍。

  2. @omarsar076

    阿里 Qwen 团队开放 Qwen3.8-Flash 权重,该模型为多模态 MoE,总参数 125B、每 token 激活 6B,并带 51B N-gram 嵌入,官方称其为 Qwen4 架构的早期预览。生产版本将上线 QwenCloud API,输入 $0.16/1M tokens、输出 $0.47/1M tokens,官方还给出 DeepSWE 1.1 58.7、SWE-bench Pro 62.5 等成绩。作者 @omarsar0 认为随附的技术报告比发布本身更值得读。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:发布信息列出了 Qwen3.8-Flash 的参数量、激活规模与定价,可作为判断高效多模态 MoE 路线的具体参照。

  3. @AYi_AInotes76

    Qwen 团队开源 Qwen3.8-Flash,总参数 125B 加 51B N-gram 嵌入,每 token 仅激活 6B,训练成本为 Qwen3.7-Plus 的 1/9。

    引用@Alibaba_Qwen@Alibaba_Qwen

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://t.co/M5hYypFLgJ - Technical Report: https://t.co/IF0gObIkQO - Hugging Face: https://t.co/6ow8QVAABt - ModelScope: https://t.co/tDOn2jNuFG

    推荐理由:Qwen3.8-Flash 以 125B 总参数、6B 激活和 1/9 训练成本给出开源 MoE 的效率样本,可对照其基准数据看架构取舍。

  4. @testingcatalog76

    阿里发布 Qwen3.8 Flash,一款 125B 参数的多模态 MoE 模型,原生上下文 262K,可通过 YaRN 扩展至 1M。QwenCloud API 定价为每 1M 输入 tokens 0.16 美元、每 1M 输出 tokens 0.47 美元。该模型基于新架构,是 Qwen4 所用架构的前身,在 DeepSWE 1.1 得 58.7 分、SWE-bench Pro 得 62.5 分。

    引用@Alibaba_Qwen@Alibaba_Qwen

    Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.

    推荐理由:原文给出上下文长度、API 定价与多项编码基准分数,读者可据此对比同表内 DeepSeek 与 Claude 模型的定位。

  5. 机器之心 · 微信公众号77

    阿里发布 Qwen3.8-Flash,同步开源 Qwen3.8-Flash-Next

    阿里发布 Qwen3.8-Flash,并在 Hugging Face 与 ModelScope 开放同一模型的 Qwen3.8-Flash-Next 权重,主模型 125B 参数、每 token 仅激活 6B,千问 AI 平台定价为每百万 token 输入 1 元、输出 3 元。

    推荐理由:文章拆解了 Qwen3.8-Flash 在注意力、残差与嵌入上的四处架构改动,并把它放进每任务成本的行业对比框架里。

  6. @kimmonismus80

    Qwen3.8-Flash-Next 发布,采用 125B MoE 参数加 51B N-gram embeddings,每 token 仅激活 6B 参数。

    引用@kimmonismus@kimmonismus

    Qwen 3.8 Flash-Next official released: A 6B-active open model just beat Claude Opus 4.6 Max across 8 of 9 comparable benchmarks! Qwen3.8-Flash-Next is a highly sparse MoE: • 125B model parameters • 51B additional n-gram embeddings • Only 6B parameters active per token It scores: • 62.5 SWE-bench Pro • 81.0 SWE-bench Multilingual • 73.9 CoworkBench • 55.7 JobBench • 73.5 Toolathlon • 81.3 IFBench • 91.7 GPQA Diamond • 91.9 LiveCodeBench It also outperforms Qwen3.8-27B and DeepSeek-V4-Flash across most of the table. Super cool release!!

    推荐理由:原文给出四项架构改动与 1/9 训练成本的对比,读者可以了解高稀疏 MoE 如何压低单 token 计算量。

  7. Qwen Blog69

    Qwen3.8-Flash-Next 开源,多模态 MoE 架构预览 Qwen4

    千问团队开源 Qwen3.8-Flash-Next 权重,这是一个多模态 MoE 模型,也是 Qwen4 所用架构的早期预览。文中称其角色类似 Qwen3-Next 之于 Qwen3.5,当时的混合 Gated DeltaNet + Gated Attention 设计已用于 Qwen3.5 至 Qwen3.8 系列。

    推荐理由:官方开源权重并定位为 Qwen4 架构预览,读者可据此追踪千问后续系列的架构走向。

8月25日周二
  1. @alibaba_cloud69

    Wan 3.0 已在 Runway 上线,可生成视频和音频,并支持输入多张图像、视频和音频作为参考,官方称这让每次生成获得更多控制。该模型还可通过 Model Studio 和 Qwen Cloud 获取 API 访问。

    引用@runwayml@runwayml

    WAN 3.0 is now on Runway. Generate state-of-the-art video and audio with multiple image, video and audio reference inputs. Try it now at the link below. https://t.co/VzTkPZR3cz

    推荐理由:Wan 3.0 上线 Runway 并给出 API 入口,读者可了解其多模态参考输入带来的生成控制方式。