Martian 发布 AI Frontier 交互式仪表盘,从质量、真实工作负载成本和重复运行一致性对比 44 个 LLM。仪表盘分为五个视图,包括单模型档案、两两对比、报价与实际成本、重复尝试一致性以及底层方法论。其中 Model Reliability 衡量模型在重复尝试中能否以相同方式解决同类问题,分数大致在 0.74 到 0.96 之间,没有单一模型或模型家族占据绝对领先。
Martian has released AI Frontier, an interactive dashboard that measures how 44 LLMs compare on quality, cost in real workloads, and how consistently each one can solve the same problem across repeated runs.
The dashboard is split into five views: single-model profiles, head-to-head comparisons, quoted price vs. measured cost, consistency across repeat attempts, and the underlying methodology.
Model Reliability tracks how often a model can solve the same kind of problem the same way across repeated attempts, with scores running from roughly 0.74 to 0.96.
No single model or family has an absolute lead 👀
来源:@testingcatalog · x.com