跳到正文
@Alibaba_Qwen· @Alibaba_Qwen · X·· 2026-08-26AI 评分37
AI 导读

通义千问公布新模型架构的四项核心升级:注意力采用 GDN + QSA 混合,Gated DeltaNet 压缩历史、Qwen Sparse Attention 用轻量索引器做微块上下文选择,降低长序列注意力成本;残差改为 Gated Residual,将残差流扩展至 4 条分支并引入动态读写门控,增强跨层信息流并显著提升训练稳定性。

正文

Model Architecture

Four core upgrades for maximum capability, efficiency, capacity, and stability:

- Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences.
- Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability.
- Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching.
- Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.

来源:@Alibaba_Qwen · x.com