跳到正文
原文
@testingcatalog· @testingcatalog · X·· 2026-08-26精选AI 评分76
AI 导读

阿里发布 Qwen3.8 Flash,一款 125B 参数的多模态 MoE 模型,原生上下文 262K,可通过 YaRN 扩展至 1M。QwenCloud API 定价为每 1M 输入 tokens 0.16 美元、每 1M 输出 tokens 0.47 美元。该模型基于新架构,是 Qwen4 所用架构的前身,在 DeepSWE 1.1 得 58.7 分、SWE-bench Pro 得 62.5 分。

推荐理由

原文给出上下文长度、API 定价与多项编码基准分数,读者可据此对比同表内 DeepSeek 与 Claude 模型的定位。

正文

Alibaba released Qwen3.8 Flash, a new multimodal 125B-parameter MoE model.

> 262K native context, extensible to 1M with YaRN.
> Pricing on QwenCloud API - $ 0.16/1M input tokens and $ 0.47/1M output tokens.
> Built on top of the new architecture, serving as a precursor to the architecture used in Qwen4.
> Qwen3.8 Flash scores 58.7 on DeepSWE 1.1 and 62.5 on SWE-bench Pro.

引用@Alibaba_Qwen@Alibaba_Qwen
Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.
在 X 查看被引用的帖子

来源:@testingcatalog · x.com