AI 导读
atomic_chat_hq 对 Qwen 3.8 27B 的量化测试显示,更高量化精度并不自动带来更好结果,Q8 在部分 voxel 任务上仅与 Q4 持平甚至被反超。
正文
Super interesting test by @atomic_chat_hq
Shows that more quantization precision does not automatically produce the better result.
Here, Qwen 3.8 27B Q8 (8-bit quantized version of the model) sometimes matched or beat Q4 (4-bit quantized) on the same voxel tasks.
They finally recommend AD-Q5_K_M as their preferred quant, since it runs on a 32GB MacBook Air with 32K context and retains 97.3% next-token agreement with BF16.
For the model weights alone:
Q8_0: 28.9GB
AD-Q4_K_M: 17.1GB
So Q4 saves 11.8GB of RAM/VRAM, or about 40.8% versus Q8.
来源:@rohanpaul_ai · x.com