跳到正文
@TencentHunyuan· @TencentHunyuan · X·· 2026-08-29AI 评分46
AI 导读

腾讯混元将 Hy4-preview 从 1.5TB 压缩至约 200GiB GGUF,提出 MIX-STQ1_0 混合量化方案,按校准数据为每层选择位宽,部分低至 1.31-bit STQ1_0、部分高至 2.06-bit IQ2_XXS。

正文

We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well !

Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error.

Accuracy barely moves vs BF16
📊 MCP Atlas 83.7→83.2
📊 SWE-Bench multi 82.9→81.3
📊 MRCR 81.3→81.1
📊 IFBench 73.5→72.5

See the details on HF : AngelSlim/Hy4-preview-GGUF

Weights & low-bit GGUFs 👇
https://t.co/9uM9NT9Wem

#LLM #Quantization #llamacpp #Hy

引用@TencentHunyuan@TencentHunyuan
🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog:https://t.co/rbl1IWRk3C HuggingFace:https://t.co/mE9wevH5XR Github:https://t.co/pyl9zckpoL https://t.co/4iW6gSuZKr
在 X 查看被引用的帖子

来源:@TencentHunyuan · x.com