跳到正文
Georgi Gerganov· @ggerganov · X·· 2 小时前AI 评分59
AI 导读

llama.cpp 作者 Georgi Gerganov 介绍,llama.cpp 可通过 ggml RPC 后端在异构设备间分配推理。他称这是一个进阶设置,未来会逐步让普通用户更容易使用。引用中提到在 RTX 6000 GPU 与 M5 笔记本上以 mxfp4 原生权重跑 MiMo 2.6 Flash,经 10 GbE 达到 40 tokens/sec,该功能在 llama.cpp 中开箱即用。

正文

llama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend

It's an advanced setting but I think with time we'll make it more accessible to regular users.

引用Pedro Cuenca@pcuenq
It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.
在 X 查看被引用的帖子

来源:Georgi Gerganov · x.com