跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 29 天前AI 评分37
AI 导读

在 GDPval-AA v2 上(该基准以 1,000 的人类基线测试模型在真实工作任务中的表现),MiniCPM5-2B 达到 831 的 Elo,领先 Ling 3.0 Tiny(718)约 110 分,领先 Granite 4.2 8B(647)约 180 分。这一规模的模型通常远低于此,LFM2.5-2.6B 为 204,Gemma 4 E4B(Reasoning)为 178。

正文

On GDPval-AA v2, which tests models on real-world work tasks against a human baseline of 1,000, MiniCPM5-2B reaches an Elo of 831, ~110 points ahead of Ling 3.0 Tiny (718) and ~180 ahead of Granite 4.2 8B (647). Models at this scale usually sit far lower, with LFM2.5-2.6B at 204 and Gemma 4 E4B (Reasoning) at 178.

来源:@ArtificialAnlys · x.com