AI 导读
llama.cpp 作者 Georgi Gerganov 介绍,llama.cpp 可通过 ggml RPC 后端在异构设备间分配推理。他称这是一个进阶设置,未来会逐步让普通用户更容易使用。引用中提到在 RTX 6000 GPU 与 M5 笔记本上以 mxfp4 原生权重跑 MiMo 2.6 Flash,经 10 GbE 达到 40 tokens/sec,该功能在 llama.cpp 中开箱即用。
正文
llama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend
It's an advanced setting but I think with time we'll make it more accessible to regular users.
It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.在 X 查看被引用的帖子
来源:Georgi Gerganov · x.com