跳到正文
原文
Artificial Analysis· @ArtificialAnlys · X·· 1 天前AI 评分44
AI 导读

在衡量从漏洞发现到修补全流程网络安全能力的 CyberGym-E2E-AA 基准上,部分前沿智能模型因安全限制在 85%+ 的任务上被拦截无法响应。GPT-6 Luna 或 MiMo-V2.6-Pro 可在 100 万行以上代码库中执行约 100 次漏洞挖掘,成本约 20 美元,单任务成本比次优模型 Grok 4.7 便宜最多 100 倍。

正文

Restricting offensive without blocking defensive is difficult, as finding and proving a vulnerability requires the same steps whether the goal is to exploit it or patch it

On CyberGym-E2E-AA, a benchmark that measures cyber defense capabilities from discovery to patching on memory-safety tasks, some frontier intelligence models are safety blocked from responding on 85%+ of tasks.

The good news is that some of the most capable models are also the most cost effective. With GPT-6 Luna or MiMo-V2.6-Pro, you can run ~100 bug hunts in a 1M+ line codebase for ~$20 - up to 100x cheaper per task than the next most capable model Grok 4.7.

来源:Artificial Analysis · x.com