AI 导读
测试框架的差异是更有意思的数字。 例如,GPT-5.6 Luna 在一个框架上得分 33.6%,在另一个框架上得分 44.9%。同一个模型,同样的 107 个任务,相差 11 分。这比榜单上第二名和第八名之间的差距还大。 在长周期商业任务上,模型周围的脚手架对结果的影响几乎和模型本身一样大。 到那时,人类仍然做出关键决策,而智能体执行一部分任务。
正文
The harness spread is the more interesting number.
For instance, GPT-5.6 Luna scores 33.6% on one harness and 44.9% on another. Same model, same 107 tasks, an 11-point swing. That is wider than the gap between second and eighth place on the board.
On long-horizon commerce tasks, the scaffolding around a model can move the result about as much as the model does.
At that point, humans still make the key decision, while agents execute a share of the tasks.
来源:@testingcatalog · x.com