跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 27 天前AI 评分51
AI 导读

Artificial Analysis 上线 Model Release 页面,可在同一处对比同一模型各推理强度变体的智能指数、每任务成本、输出速度与延迟。每个页面还提供与同类模型各推理强度的并排比较,并给出金融与会计、法律、医疗与健康、战略与运营、工程、经济六类 Capability Index 分数。

正文

You can now easily compare the intelligence, cost, and speed on different effort levels for your preferred models with the new Artificial Analysis Model Release pages

How a model performs across intelligence and cost is heavily influenced by its configured reasoning and effort level. Frontier models are now being released with up to six different effort levels, meaning the same model weights can produce significantly different performance and cost profiles. Model Release pages make it easy to compare these configurations across all of our standard charts in one place

Each release page features:
➤ Intelligence, cost per task, output speed, and latency across every effort variant
➤ Side by side comparison of effort levels against those of similar models
➤ Artificial Analysis Capability Index scores on each effort level across Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics

For example, on Intelligence vs Time per Task, GPT-6 Astra ranges from 46–53 on the Artificial Analysis Intelligence Index and 1.6–8.2 minutes per task, while Claude Fable 5.1 spans a similar intelligence range (47–53) on a wider time per task range (4.2–12.2 minutes)

来源:@ArtificialAnlys · x.com