跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-20AI 评分30
AI 导读

Apodex 推出 TRACES,号称全球首个用于衡量"发现式 AI"的基准,主张把 AI 发现能力当作完整调查过程来评估,而非只看单一最终答案。该基准关注系统在证据、工具与验证条件下调查重要问题的能力,可区分高分结果与产生该结果的过程质量,包括错误是否被修复、结论是否有据可依。

正文

This is a good example of why final-answer accuracy can hide bad agent behaviour.

Most AI benchmarks measure whether a model can reach a known answer.

Apodex introduced TRACES 🧭, the world's first benchmark for measuring discoverative AI.

A shift from benchmarking models on solved problems to evaluating systems that can investigate consequential problems under evidence, tools, and verification.

TRACES says AI discovery should be evaluated as an entire investigation, not as a single final answer.

This can separate a high-scoring outcome from the quality of the process that produced it, including whether errors were repaired and claims were grounded.

来源:@rohanpaul_ai · x.com