跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-08-25AI 评分34
AI 导读

Artificial Analysis 公布多项小模型单项评测结果:Qwen3.5 9B(非推理)以 77% 拿下 BFCL、79% 拿下 GPQA Diamond,Falcon-H1R-7B 以 97% 拿下 MATH-500,LFM2.5-2.6B 以 59% 拿下 IFBench 并在 AA-Omniscience 上实现 79% 非幻觉率。

正文

In the individual evaluations: Qwen3.5 9B (Non-reasoning) takes BFCL at 77% and GPQA Diamond at 79%, Falcon-H1R-7B takes MATH-500 at 97%, LFM2.5-2.6B takes IFBench at 59% and hallucinates far less than the other leading models on AA-Omniscience (79% non-hallucination), Ornith-1.0-9B recalls the most facts (15% accuracy), and G9v3-3B has the highest non-hallucination rate of any model that attempts answers (89%). Nanbeige4.2-3B wins none of the evaluations outright but is top-five on BFCL, GPQA Diamond and MATH-500 - which is how it ties for first overall

来源:@ArtificialAnlys · x.com