Martian 发布 AI Frontier 仪表盘,可按任务、质量、实际成本与可靠性对比 LLM,并评估组合多模型是否效果更好。在 16 项 benchmark 上,其 oracle 路由在同等成本下将平均错误率降低 54%,或在匹配各 benchmark 最佳模型质量的同时把 API 成本降低 85%。
Picking one "best model" is starting to look like the wrong unit of optimization.
@withmartian just released AI Frontier, a dashboard for comparing LLMs by task, quality, actual cost, and reliability, then seeing whether combining different models gives better results.
Across 16 benchmarks, Martian’s oracle routing cut average error 54% at matched cost, or matched each benchmark’s top-model quality at 85% lower API cost.
The idea is that there may be no single "best" LLM: Different models perform better on different workloads, so AI Frontier shows when choosing or combining models can produce a better mix of quality, cost, and reliability than using one model for everything.
Output length, reasoning behavior, retries, and reliability all affect what you eventually pay to get a usable answer.
Martian is also measuring how consistently models solve the same kinds of problems, which makes the cost picture more useful. A cheap model that needs several attempts may not be cheap at the system level.
Model economics should probably be measured as cost per acceptable result, not simply dollars per million tokens.
来源:@rohanpaul_ai · x.com