在标准 16K 上下文限制下,Nanbeige4.2-3B(Reasoning)与 LFM2.5-2.6B(Reasoning)以平均评测分 63 并列第一,领先 Ornith-1.0-9B(62)、Qwen3.5 9B(Reasoning,61)、Ornith-1.5-9B(61)、Gemma 4 E4B(60)和 Qwen3.5 9B(Non-reasoning,60)。
At the standard 16K context limit, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tie for the top average evaluation score at 63, ahead of Ornith-1.0-9B (Reasoning, 62), Qwen3.5 9B (Reasoning, 61), Ornith-1.5-9B (Reasoning, 61), Gemma 4 E4B (Reasoning, 60) and Qwen3.5 9B (Non-reasoning, 60). Ornith-1.5-9B slips behind its older 1.0 sibling purely on a weaker instruction following performance in IFBench
When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66, followed by Nanbeige4.2-3B at 65, and both Qwen3.5 9B (Reasoning) at 64 and LFM2.5-2.6B at 64
来源:@ArtificialAnlys · x.com