Apodex AI 发布 TRACES 基准,用于测试 AI 系统在正确答案未知的困难科学问题上的表现,并分别衡量 Tools、Repair、Alternatives、Coherence、Evidence、Scope 六项能力。
A benchmark for difficult scientific discovery from @tianqiao_chen now shows GPT-6 Astra improving across all six capabilities.
But fixing its own mistakes improved the least, Repair improved far less than the other dimensions.
For AI agents, a correct final answer can hide a weak investigation.
An agent might reach the right result after ignoring contradictory feedback. That behavior could fail on the next task, even though the current answer passes its check.
@Apodex_AI released TRACES, a new benchmark for testing AI systems on difficult scientific problems where the correct answer may not already be known.
TRACES separately measures Tools, Repair, Alternatives, Coherence, Evidence, and Scope instead of reducing discovery capability to one score.
Instead of scoring only the final answer, TRACES also evaluates how the AI works through the problem, including its tool use, error correction, evidence, and reasoning process.
来源:@rohanpaul_ai · x.com