AI 导读
🧠 非对称架构。更强智能,更低成本。 🔹 552B 参数 MoE。 🔹 全新因果编码器–解码器架构:输入仅 8B 激活参数,输出 16B。 🔹 全新预训练方法 + 更大规模 RL 后训练,基准测试结果超越旗舰模型,包括 DeepSeek-V4-Pro。 2/6
正文
🧠 Asymmetric architecture. More intelligence, less cost.
🔹 552B-parameter MoE.
🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
2/6
来源:@deepseek_ai · x.com