AI 导读
这种效率很大程度上来自 Mixture-of-Value Attention,即 MoVA,它把 Mixture-of-Experts 路由扩展到注意力中,而不是只在 feed-forward 层内应用稀疏性。 所以这个 36B 模型有 36B 总参数,但每个 token 只有约 4B 激活,同时仍接近稠密的 32B 模型。 而且因为这些模型是在相同条件下训练的,它们为研究者提供了一个相当干净的稠密 vs 稀疏对比,而不是比较不相关的模型。
正文
A lot of that efficiency comes from Mixture-of-Value Attention, or MoVA, which extends Mixture-of-Experts routing into attention instead of applying sparsity only inside the feed-forward layers.
So the 36B model has 36B total parameters but only about 4B active for each token, while still landing close to the dense 32B model.
And because those models were trained under the same conditions, they give researchers a fairly clean dense-vs-sparse comparison instead of comparing unrelated models.
来源:@rohanpaul_ai · x.com