跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 29 天前AI 评分40
AI 导读

Artificial Analysis 将评测基准从 Terminal-Bench 2.1 升级至 4.0,覆盖软件、机器学习、科学、运维、安全、硬件与媒体等终端复杂任务,重新校准算力与时间配额并改进任务说明、环境与验证,全部 66 项任务各跑三次取平均 pass@1,统一使用 mini-SWE-agent 测试框架。

正文

We are upgrading from Terminal-Bench 2.1 to 4.0, which features harder tasks from a terminal. Terminal-Bench 4.0 tests whether an agent can complete complex work through the terminal, across software, machine learning, science, operations, security, hardware, and media. The update recalibrates compute and time allowances and improves task instructions, environments, and verification. We run all 66 tasks three times and report average pass@1. We use mini-SWE-agent, a minimal, model-agnostic harness, for all runs

GPT-6 Astra (max) scores 59.1%, compared with 52.0% for Claude Fable 5.1 (max with fallback) and 49.0% for Claude Opus 5 (max). Astra is 19 percentage points ahead of GPT-5.6 Sol (max), which scores 39.9%

来源:@ArtificialAnlys · x.com