跳到正文
@testingcatalog· @testingcatalog · X·· 2026-08-29AI 评分42
AI 导读

测试框架的差异是更有意思的数字。 例如,GPT-5.6 Luna 在一个框架上得分 33.6%,在另一个框架上得分 44.9%。同一个模型,同样的 107 个任务,相差 11 分。这比榜单上第二名和第八名之间的差距还大。 在长周期商业任务上,模型周围的脚手架对结果的影响几乎和模型本身一样大。 到那时,人类仍然做出关键决策,而智能体执行一部分任务。

正文

The harness spread is the more interesting number.

For instance, GPT-5.6 Luna scores 33.6% on one harness and 44.9% on another. Same model, same 107 tasks, an 11-point swing. That is wider than the gap between second and eighth place on the board.

On long-horizon commerce tasks, the scaffolding around a model can move the result about as much as the model does.

At that point, humans still make the key decision, while agents execute a share of the tasks.

来源:@testingcatalog · x.com