跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-08-25AI 评分31
AI 导读

在标准 16K 上下文限制下,Nanbeige4.2-3B(Reasoning)与 LFM2.5-2.6B(Reasoning)以平均评测分 63 并列第一,领先 Ornith-1.0-9B(62)、Qwen3.5 9B(Reasoning,61)、Ornith-1.5-9B(61)、Gemma 4 E4B(60)和 Qwen3.5 9B(Non-reasoning,60)。

正文

At the standard 16K context limit, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tie for the top average evaluation score at 63, ahead of Ornith-1.0-9B (Reasoning, 62), Qwen3.5 9B (Reasoning, 61), Ornith-1.5-9B (Reasoning, 61), Gemma 4 E4B (Reasoning, 60) and Qwen3.5 9B (Non-reasoning, 60). Ornith-1.5-9B slips behind its older 1.0 sibling purely on a weaker instruction following performance in IFBench

When the context limit is raised to 64K (represented by dots in the image), Ling 3.0 Tiny takes the top spot at 66, followed by Nanbeige4.2-3B at 65, and both Qwen3.5 9B (Reasoning) at 64 and LFM2.5-2.6B at 64

来源:@ArtificialAnlys · x.com