通义千问公布新模型架构的四项核心升级:注意力采用 GDN + QSA 混合,Gated DeltaNet 压缩历史、Qwen Sparse Attention 用轻量索引器做微块上下文选择,降低长序列注意力成本;残差改为 Gated Residual,将残差流扩展至 4 条分支并引入动态读写门控,增强跨层信息流并显著提升训练稳定性。
Model Architecture
Four core upgrades for maximum capability, efficiency, capacity, and stability:
- Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences.
- Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability.
- Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching.
- Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture.
来源:@Alibaba_Qwen · x.com