@berryxia· @berryxia · X·· 2026-05-14精选AI 评分72
AI 导读
Unsloth 创始人 Daniel Han 发布 Qwen3.6 实验版 MTP GGUF,27B 模型单 GPU 生成速度达 140 tokens/s,35B-A3B 版本达 220 tokens/s。
推荐理由
原文给出单 GPU 实测速度、加速比与 draft tokens 甜点,可据此判断本地 30B 级模型的部署空间。
正文
我靠,肉眼都跟不上这个速度了!
Daniel Han,UnslothAI创始人,YC S24,之前在NVIDIA做ML,刚刚把Qwen3.6的实验MTP GGUF放出来了。
27B模型单GPU直接跑到140 tokens/s。
35B-A3B版本更猛,冲到220 tokens/s。
比原版GGUF快超过1.4倍,精度零损失。
他们测了半天,发现draft tokens设成2就是甜点,再往上接受率暴跌,实际速度反而掉下去。
我看完那张benchmark曲线图,最大的感受是,本地大模型的性能天花板又被狠狠顶高了一截。
以前总觉得30B+模型本地跑太慢,现在MTP投机解码直接把消费级显卡的潜力榨干了。
如果你在玩llama.cpp、跑本地Agent或者日常coding,这波更新必须马上试。
本地AI越来越不像“妥协版”了。
We released experimental MTP Qwen3.6 Unsloth GGUFs! Qwen3.6 27B MTP now runs at 140 tokens/s. Qwen3.6 35B-A3B MTP gets 220 tokens/s generation on a single GPU. Qwen3.6 27B and 35B-A3B have >1.4x speed-up over the original GGUFs without any change in accuracy. Guide + GGUFs + Benchmarks: unsloth.ai/docs/models/qwen3… In terms of average speedup, we see a 1.4x for dense models at draft tokens = 2 and for the MoE around 1.15 to 1.2x. We do not recommend more than 2 draft tokens because the acceptance rate drops precipitously from 83% to 50% with 4 draft tokens, and the forward passes for MTP become less beneficial. Use `--spec-type mtp --spec-draft-n-max 2` Thanks to Aman for github.com/ggml-org/llama.cp…!在 X 查看被引用的帖子
来源:@berryxia · x.com