跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 25 天前AI 评分62
AI 导读

K2 Horizon 推出 Uno Diffusion,一个在冻结原自回归模型的前提下学习并行生成 token 块的 LoRA 适配器,因此无需单独的草稿模型,也不必迁移基座模型。IFM 称其推理速度约提升 3 倍且模型质量无损。该方案与整套训练、对话、工具调用和部署栈共用,对 vLLM、SGLang 和 Ollama 提供 day-zero 支持,覆盖 NVIDIA、AMD 和 Cerebras 硬件。

正文

K2 Horizon also ships Uno Diffusion, a LoRA adapter that leaves the original autoregressive model frozen while learning to generate blocks of tokens in parallel.

That means no separate draft model and no base-model migration, while IFM reports roughly 3x faster inference with no loss in model quality.

And the whole fleet shares the same basic training, chat, tool-calling and deployment stack, with day-zero support for vLLM, SGLang and Ollama across NVIDIA, AMD and Cerebras hardware.

来源:@rohanpaul_ai · x.com