AI 导读
面壁智能发布 MiniCPM5-2B,这是一款 2.6B 参数的稠密推理模型,仅支持文本输入输出,采用 Apache 2.0 许可,上下文窗口 131k tokens。在 Artificial Analysis Intelligence Index v4.2 上得分 15,为 4B 以下开源权重模型中的最高分,比 Granite 4.2 3B 高 4 分。官方称其优化方向是 2.6B 规模下的智能体表现与 token 效率,即端侧的高智能密度,该模型每个 Intelligence Index 任务消耗 19k 输出 token。
正文
Thank you @ArtificialAnlys for the independent evaluation. On Intelligence Index v4.2, MiniCPM5-2B scores 15, the highest of any open weights model under 4B total parameters.
Agentic performance + token efficiency at 2.6B is exactly what we optimized for — high intelligence density on the edge.
🤗 https://t.co/TNUu629lXn
💻 https://t.co/zaffsLsx5m
OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model under 4B total parameters OpenBMB (@OpenBMB) is the open-source AI group behind the MiniCPM series of efficient small models. MiniCPM5-2B is a 2.6B parameter dense reasoning model with text input and output, released under Apache 2.0. Scoring 15 on the Intelligence Index, MiniCPM5-2B sits one point behind Ling 3.0 Tiny (16), which has ~3x the total parameters. Among open weights models under 4B total parameters, the next best score is Granite 4.2 3B (11). Key results: ➤ The highest Intelligence Index of any open weights model under 4B total parameters, setting a new Pareto-optimal point on Intelligence vs. Total Parameters: Its score of 15 is 4 points clear of Granite 4.2 3B (11). With 2.6B total parameters, it is 1 point ahead of Qwen3.5 4B (Reasoning, 14, estimated) with 44% fewer parameters, and level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size. As a dense model, its size advantage is in memory footprint rather than active-parameter compute. ➤ Strong agentic performance at this size: Its GDPval-AA v2 Elo of 831 leads <4B models, and on τ³-Banking it is joint-first with Ling 3.0 Tiny at 21%, compared to 8% for the next best model, Granite 4.2 8B. On AA-Briefcase, it placed second among the measured models in the comparison set with an Elo of 438, above Granite 4.2 8B (324) and just below Ling 3.0 Tiny (485). ➤ Knowledge, coding and long context are where it gives ground: MiniCPM5-2B places 7th in the set on Humanity's Last Exam (9%, behind Gemma 4 12B (Reasoning) at 16%), 8th on Terminal-Bench v2.1 (9%, behind Qwen3.5 9B (Reasoning) at 29%) and scores 0% on CritPt. On SciCode it is second of the five measured models at 26%, behind Granite 4.2 8B (31%). On AA-LCR v1.1 it scores 59%, 5th in the set, one point behind Ling 3.0 Tiny (60%). On GDP.pdf, our new professional document reasoning evaluation, it passes 1% of tasks outright, behind gpt-oss-20b (high) at 2%. ➤ Its AA-Omniscience score of -12 is earned by abstaining from answering rather than accuracy: MiniCPM5-2B attempts only 29% of AA-Omniscience questions, giving it a Non-Hallucination Rate of 78%. Its accuracy of 8% is a point below Ling 3.0 Tiny (9%) and half that of Qwen3.5 9B (Reasoning, 16%). Peers that attempt far more questions are penalized heavily, with Qwen3.5 9B (Reasoning) at -53 and gpt-oss-20b (high) at -63. ➤ It is token-efficient for a reasoning model: MiniCPM5-2B used 19k output tokens per Intelligence Index task, joint-lowest in the comparison model set with Granite 4.2 3B (19k). Ling 3.0 Tiny spends 56k, roughly 3x as many, for 1 more index point. Additional model details: ➤ Parameters: 2.6B (dense) ➤ Context window: 131k tokens ➤ Input modalities: Text only ➤ License: Apache 2.0在 X 查看被引用的帖子
来源:@OpenBMB · x.com