跳到正文
@SemiAnalysis_· @SemiAnalysis_ · X·· 2026-08-27AI 评分51
AI 导读

阿里 Qwen 发布 Qwen3.8-Flash-Next,SemiAnalysis 指出它采用了与即将推出的 Qwen4 相同的架构创新,包括 51-billion-param N-gram Embedding(可用极少额外算力查表。

正文

Congrats to @Alibaba_Qwen on the release of Qwen3.8-Flash-Next, using the same architecture innovations as their upcoming Qwen4 model! Such innovations include:

🟠 51-billion-param N-gram Embedding to look up a table with very little extra computation, which means the embedding table can be offloaded to slower & less expensive tiers of DRAM
🟠 Gated Residual (GR): it seems like a lot of Chinese labs are now innovating on the res connections, like Kimi's AttentionRes and DeepSeek's mHC
🟠 Qwen Sparse Attention (QSA): lightning indexer to select context at micro-block granularity

Glad to see great Chinese open innovations along with end-to-end model weights to show these innovations can compose well together!

来源:@SemiAnalysis_ · x.com