跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 22 天前AI 评分61
AI 导读

耶鲁大学等机构的研究者让物理学家重新核查模型在 6 个物理基准上的失败案例,发现 250 个被拒案例中只有 12 个是模型真正出错,其余 238 个来自题目、参考答案或评分器问题。经专家复核,GPT-5.6-Sol 的 HLE-Physics 分数从 47.3% 升至 78.7%,但作者智能体仍未完全解决任何尝试的开放理论物理问题。

正文

New Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics benchmarks.

And many apparent failures come from bad questions, wrong reference answers, and brittle graders explain most audited physics failures, making benchmark quality the new bottleneck for measuring frontier models.

The researchers had physicists re-check model failures across 6 popular physics benchmarks instead of trusting the original scores.

In 250 rejected cases from 4 audited benchmark subsets, only 12 were actual model mistakes.

The other 238 came from bad questions, wrong reference answers, or graders rejecting correct answers.

After expert review, GPT-5.6-Sol’s measured HLE-Physics score rose from 47.3% to 78.7%.

So a low physics benchmark score can badly underestimate what a frontier model can actually solve.

But that does not mean these models can reliably do physics research: the authors’ agents still failed to fully solve any of the open theoretical-physics problems they tried.

but it does mean that do not treat benchmark scores as clean ground truth anymore for AI's Physics capability.

来源:@rohanpaul_ai · x.com