跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-25AI 评分35
AI 导读

atomic_chat_hq 对 Qwen 3.8 27B 的量化测试显示,更高量化精度并不自动带来更好结果,Q8 在部分 voxel 任务上仅与 Q4 持平甚至被反超。

正文

Super interesting test by @atomic_chat_hq

Shows that more quantization precision does not automatically produce the better result.

Here, Qwen 3.8 27B Q8 (8-bit quantized version of the model) sometimes matched or beat Q4 (4-bit quantized) on the same voxel tasks.

They finally recommend AD-Q5_K_M as their preferred quant, since it runs on a 32GB MacBook Air with 32K context and retains 97.3% next-token agreement with BF16.

For the model weights alone:

Q8_0: 28.9GB
AD-Q4_K_M: 17.1GB

So Q4 saves 11.8GB of RAM/VRAM, or about 40.8% versus Q8.

来源:@rohanpaul_ai · x.com