Martian 发布 AI Frontier,用每个数据点 10 次运行、覆盖 16 个基准的方式衡量模型在相同提示词下的可靠性,并给出榜单。榜单显示 Qwen3.7 Max 为 96.1%、Claude Opus 4.6 为 94.4%、GPT-5.5 为 93.5%,这里的可靠性指答对或答错的一致性,而非原始准确率。
A model can look strong on average and still be unpredictable on the same prompt.
Martian’s new AI Frontier measures that gap with 10 runs per datapoint across 16 benchmarks. Its current reliability table puts Qwen3.7 Max at 96.1%, Claude Opus 4.6 at 94.4%, and GPT-5.5 at 93.5%.
Reliability here means consistency in being right or wrong, not raw accuracy. In an agent workflow, one unstable step can derail everything that follows.
The paper behind the dashboard found a 54% error reduction at matched cost using oracle routing across models versus each benchmark’s best single model.
来源:@kimmonismus · x.com