AI 导读
在 GDPval-AA v2 上(该基准以 1,000 的人类基线测试模型在真实工作任务中的表现),MiniCPM5-2B 达到 831 的 Elo,领先 Ling 3.0 Tiny(718)约 110 分,领先 Granite 4.2 8B(647)约 180 分。这一规模的模型通常远低于此,LFM2.5-2.6B 为 204,Gemma 4 E4B(Reasoning)为 178。
正文
On GDPval-AA v2, which tests models on real-world work tasks against a human baseline of 1,000, MiniCPM5-2B reaches an Elo of 831, ~110 points ahead of Ling 3.0 Tiny (718) and ~180 ahead of Granite 4.2 8B (647). Models at this scale usually sit far lower, with LFM2.5-2.6B at 204 and Gemma 4 E4B (Reasoning) at 178.
来源:@ArtificialAnlys · x.com