跳到正文
@omarsar0· @omarsar0 · X·· 2026-09-04AI 评分41
AI 导读

先说这个基准测试为什么存在。 HLE、FrontierMath、MMLU 和 BrowseComp 都是高难度测试。它们每一个都有标准答案。 一个模型可以在这些测试上全部登顶,却仍然在开放式研究上卡住。 TRACES 测试的是更难的案例。其真实答案可能需要数月甚至数年才能确认。 Apodex 手工构建了这套题目。十位 STEM 博士花了两个月,在 16 个领域的 561 个行业中筛选,建立了一个包含 423 个高价值问题的登记库。

正文

Let's start with why this benchmark exists.

HLE, FrontierMath, MMLU, and BrowseComp are hard tests. Every one of them comes with an answer key.

A model can top all of them and still stall on open research.

TRACES tests the harder cases. The ground truth may take months or years to confirm.

Apodex built the problem set by hand. Ten STEM PhDs spent two months scouting 561 industries across 16 sectors to build a registry of 423 high-value problems.

来源:@omarsar0 · x.com