跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-08-26AI 评分47
AI 导读

Artificial Analysis 在 Coding Agent Index v1.4 中为 Terminal-Bench v2.1 引入奖励黑客修正:若某次通过的尝试被判定为奖励黑客(如直接上网抓取已公开基准数据集的答案),该次尝试记 0 分。由于 Terminal-Bench v2.1 任务未明确禁止外部搜索且运行时可访问公网,各智能体和模型的奖励黑客发生率差异很大。

正文

Introducing reward hacking score corrections to the Artificial Analysis Coding Agent Index

In v1.4 of the Artificial Analysis Coding Agent Index, we introduced reward hacking corrections to Terminal-Bench v2.1. Reward hacking is when a model successfully ‘completes’ a task without doing the work the task was meant to measure, such as deliberately fetching the solutions online for a published benchmark dataset. If a passing Terminal-Bench v2.1 attempt is found to be reward hacking, we give that attempt a zero score.

Rates vary widely by agent and by model. Unlike some evaluations, Terminal-Bench v2.1 tasks don’t explicitly instruct agents not to search for solutions externally, and the tasks run with public internet access. For a model that knows the benchmark from training data, fetching the answer is a natural but unaligned step.

来源:@ArtificialAnlys · x.com