Mistral 发布新模型 Mistral Large 4(ML4,le Chonk)预览版,总参数 1T、激活参数 49B,原生多模态,月底前发布正式版及权重。
原文整理了 Mistral Large 4 预览版在编码和网络安全基准上的具体分数和训练用量,可用作评估开源模型格局的参考。
This is a surprising great release. Le Chaton fat is real! Mistral Large 4 puts Europe back in contention on coding and cybersecurity.
Didnt expect Mistral to compete with GLM5.3!
“Le Chonk” is a natively multimodal model with 1T parameters and 49B active. The preview already delivers:
- DeepSWE v1.1: 61.7%, roughly level with GLM-5.3 at 61%. Kimi K3 remains ahead at around 68% in Mistral’s comparison.
- Artificial Analysis Cyber Index: 50, matching GLM-5.3-Flash and beating GLM-5.3’s 36.
- CyberGym-E2E-AA: 82%, ahead of MiMo-V2.6-Pro’s 79%.
Mistral says it trained the model from scratch in its own European datacenters. NVIDIA says only 4,000 Grace Blackwell Superchips powered the training! Which is impressive.
Learning: you can compete with much less compute. Which is crazy. Congrats Mistral!
Today, we are launching a preview of our new model, Mistral Large 4 (ML4), aka le Chonk 🐈. ML4 is a 1T-parameter model with 49B active parameters, trained natively with multimodal capabilities. It is at the frontier of open models, and by far the strongest open-weight model from the US or Europe. The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model. 🧵 1/n在 X 查看被引用的帖子
来源:Chubby♨️ · x.com