跳到正文
Georgi Gerganov· @ggerganov · X·· 2 小时前AI 评分43
AI 导读

如果你在使用 Qwen3.8-27B + MTP,请务必升级到 DFlash 以获得额外加速: llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-dflash --spec-draft-n-max 7 需要最新的 llama.cpp v0.6.0

正文

If you are using Qwen3.8-27B + MTP, make sure to upgrade to DFlash for extra speed:

llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-dflash --spec-draft-n-max 7

Requires the latest llama.cpp v0.6.0

引用Georgi Gerganov@ggerganov
simple: llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-mtp
在 X 查看被引用的帖子

来源:Georgi Gerganov · x.com