跳到正文
@AntLingAGI· @AntLingAGI · X·· 2026-08-22AI 评分37
AI 导读

Ling-3.0-flash 有 42 层:35 层 KDA 线性注意力层和 7 层 MLA 全注意力层。其 MoE 包含 512 个路由专家和 1 个共享专家,每个 token 激活 top 8 加 1。 这种混合设计即使在 8K 上下文下也能保持较低的注意力成本。https://t.co/h7xQrFBQHZ

正文

Ling-3.0-flash has 42 layers: 35 KDA linear attention layers and 7 MLA full attention layers. Its MoE contains 512 routed experts and one shared expert, with top 8 plus 1 active per token.
This hybrid design keeps attention cost low, even at 8K context. https://t.co/h7xQrFBQHZ

来源:@AntLingAGI · x.com