跳到正文
@deepseek_ai· @deepseek_ai · X·· 26 天前AI 评分43
AI 导读

🧠 非对称架构。更强智能,更低成本。 🔹 552B 参数 MoE。 🔹 全新因果编码器–解码器架构:输入仅 8B 激活参数,输出 16B。 🔹 全新预训练方法 + 更大规模 RL 后训练,基准测试结果超越旗舰模型,包括 DeepSeek-V4-Pro。 2/6

正文

🧠 Asymmetric architecture. More intelligence, less cost.

🔹 552B-parameter MoE.
🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

2/6

来源:@deepseek_ai · x.com