Martian 提出 Capability Frontier,一种 Pareto 式曲线,描述在每个价格点上跨模型与多次运行择优后可达的最佳分数。该曲线覆盖 21 个模型、16 个基准,涵盖编程、推理、医学、事实性、指令遵循和智能体任务。在同等成本下,它比最强单模型降低 54% 错误率,计入多次运行后达 82%,并以低 85% 的成本达到 SOTA 准确率。
Behind all of it is a Pareto-style curve that Martian calls the Capability Frontier:
> Meaning the best score reachable at each price point when the strongest answer across models and repeated runs is selected.
> Across 21 models and 16 benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks.
> It cuts error rate by 54% against the top single model at matched cost, reaching 82% once repeated runs are counted, and matches state-of-the-art accuracy at 85% lower cost.
Explore it here 👀
https://t.co/XmVnls3ExI
来源:@testingcatalog · x.com