跳到正文
原文
Karina· @karinanguyen · X·· 2 天前AI 评分52
AI 导读

Google DeepMind 发布新前沿模型 Gemini 4 Argon,面向编码、企业知识工作和网络安全防御等复杂工作流,即日起通过 Fairwind Program 向部分受信任测试者开放。作者引述其 PostTrainBench 得分 45.3%,高于 Gemini 3.1 Pro 的 21.99% 和 GPT-6 Astra 的 44.3%;评测表还显示其在自动化与智能体编码等多项基准领先,但在 FrontierSWE v2、Terminal-Bench 4.0 等项落后于对比模型。

正文

Gemini 4 reaches 45.3% on PostTrainBench, more than doubling Gemini 3.1 Pro’s 21.99% and beating GPT-6 Astra 🔥

引用Google DeepMind@GoogleDeepMind
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
在 X 查看被引用的帖子

来源:Karina · x.com