AI 导读
K2 Horizon 推出 Uno Diffusion,一个在冻结原自回归模型的前提下学习并行生成 token 块的 LoRA 适配器,因此无需单独的草稿模型,也不必迁移基座模型。IFM 称其推理速度约提升 3 倍且模型质量无损。该方案与整套训练、对话、工具调用和部署栈共用,对 vLLM、SGLang 和 Ollama 提供 day-zero 支持,覆盖 NVIDIA、AMD 和 Cerebras 硬件。
正文
K2 Horizon also ships Uno Diffusion, a LoRA adapter that leaves the original autoregressive model frozen while learning to generate blocks of tokens in parallel.
That means no separate draft model and no base-model migration, while IFM reports roughly 3x faster inference with no loss in model quality.
And the whole fleet shares the same basic training, chat, tool-calling and deployment stack, with day-zero support for vLLM, SGLang and Ollama across NVIDIA, AMD and Cerebras hardware.
来源:@rohanpaul_ai · x.com