跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-08-25AI 评分34
AI 导读

同样约 60 分,可能花 5M output tokens,也可能花 75M。Qwen3.5 9B(Reasoning)跑完一遍基准测试集要消耗 74.5M output tokens,且 29% 的生成会超出其 16K 窗口;Gemma 4 E4B(Reasoning)用 5.2M tokens 就达到相近总分,且从未触及上限。在手机上,token 消耗带来的时间、能耗和发热比笔记本、台式机或服务器更难忍受。

正文

A score of ~60 can cost 5M output tokens or 75M. Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set and runs out of its 16K window on 29% of generations; Gemma 4 E4B (Reasoning) reaches a similar overall score on 5.2M tokens and never hits the limit. On a phone, token use leads to time, energy use and heat that are less tolerable than on a laptop, desktop or server

来源:@ArtificialAnlys · x.com