AI 导读
这种扩展行为的差异可能由不同的架构选择来解释。近期开放模型采用多种方法来改善扩展行为,例如 Kimi K3 中的线性注意力、DeepSeek V4 中的稀疏注意力,或 GLM-5.3 Flash 中两者兼有。
正文
This difference in scaling behavior may be explained by different architectural choices. Recent open models use various approaches to improve scaling behavior, such as linear attention in Kimi K3, sparse attention in DeepSeek V4, or both in GLM-5.3 Flash.
来源:@EpochAIResearch · x.com