跳到正文
@berryxia· @berryxia · X·· 2026-05-15AI 评分40
AI 导读

Daily Dose of Data Science 用视觉图解释了 Transformer 与 MoE 的核心区别:MoE 将 decoder block 中的单个前馈网络拆成多个更小的专家网络,推理时只激活 top-K 个专家,参数总量更多但计算量更小、速度更快。

正文

我刚刷到 Daily Dose of Data Science 的一篇视觉解释,把 Transformer 和 Mixture of Experts(MoE)讲得特别清楚。

核心区别其实就在 decoder block:

Transformer 用的是一个大的前馈网络。

MoE 则把这个位置拆成了多个更小的“专家”网络。

推理时,MoE 只激活其中一部分专家。

参数总量虽然更多,但实际参与计算的只有一小部分,所以速度反而更快。

那模型怎么决定该激活哪些专家呢?

靠 Router。

它是一个多分类器,对每个 token 输出 softmax 分数,然后选 top-K 个专家。

但训练过程中有两个经典坑:

第一个坑是“专家过选”——一开始某个专家被选上后,它越变越强,越强越容易被选,导致其他专家几乎没机会训练。

解决办法:在 router 输出加噪声,同时把非 top-K 的 logit 直接设为 -∞,让其他专家也有训练机会。

第二个坑是“专家负载不均”——有的专家处理了太多 token,有的几乎闲着。

解决办法:给每个专家设置容量上限,超过就自动把 token 转给下一个最佳专家。

MoE 就这样用更多参数换来了更快的推理速度。

Mixtral 8x7B 和 Llama 4 都是典型的 MoE 模型。

视觉图把整个路由、专家选择、负载均衡的过程画得一目了然。

引用Daily Dose of Data Science (@DailyDoseOfDS_)@DailyDoseOfDS_
Transformer and Mixture of Experts, explained visually! Mixture of Experts (MoE) is a popular architecture that uses different experts to improve Transformer models. Transformer and MoE differ in the decoder block: - Transformer uses a feed-forward network. - MoE uses experts, which are feed-forward networks but smaller compared to those Transformer. During inference, a subset of experts are selected. This makes inference faster in MoE. Also, since the network has multiple decoder layers: - The text passes through different experts across layers. - The chosen experts also differ between tokens. But how does the model decide which experts should be ideal? The router does that. It is a multi-class classifier that produces softmax scores over experts to select the top K experts. The router is trained with the network, and it learns to select the best experts. But it isn't straightforward. There are challenges! Challenge 1) Notice this pattern at the start of training: - Say, the model selects "Expert 2" - This expert gets a bit better - It may get selected again since it's the "best" - It learns more - It gets selected again in the next iteration - It learns more, and so on! This means many experts can go under-trained due to the overselection of a few experts! We solve this in two steps: - Add noise to the feed-forward output of the router so that other experts can get higher logits. - Set all but the top K logits to -infinity. After softmax, these scores become zero. This way, other experts also get the opportunity to train. Challenge 2) Some experts may get exposed to more tokens than others, leading to under-trained experts. We prevent this by limiting the number of tokens an expert can process. If an expert reaches the limit, the token is passed to the next best expert. Overall, MoEs have more parameters to load. But a fraction of them are activated during inference. This leads to faster inference. Mixtral 8x7B and Llama 4 are two popular MoE-based LLMs. Have you used MoEs in production yet?
在 X 查看被引用的帖子

来源:@berryxia · x.com