跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-02AI 评分38
AI 导读

Apodex 发布专有模型 Apodex 1.1,在 Artificial Analysis Intelligence Index 上得分 44,在衡量真实智能体工作的 GDPval-AA v2 上达到 1,348 Elo,超过多个通用智能分数更高的模型。

正文

Apodex released Apodex 1.1, its proprietary model that reached 44 on the Artificial Analysis Intelligence Index and performs strongly on agentic tasks versus models in its tier.

on GDPval-AA v2, which measures real-world agentic work, it reaches 1,348 Elo, ahead of several models with much higher general intelligence scores.

So Apodex 1.1, its performance seems concentrated around professional and agentic tasks rather than being evenly distributed across the evaluation suite.

I increasingly think this distinction matters for model selection. A model that wins broad reasoning benchmarks is not automatically the model you want sitting inside an agentic execution loop.

For agents, the relevant question is: once you give the model a goal and tools, how often does it actually get the job done?

Apodex 1.1 from @Apodex_AI looks unusually concentrated in that direction.

来源:@rohanpaul_ai · x.com