AI 导读
llama.cpp 作者 Georgi Gerganov 宣布 GGUF 模型现在可以直接在 transformers 中运行。这项工作将 ggml 的 Metal 内核引入 transformers 生态,提升兼容性与性能;引用内容说明这些 llama.cpp 检查点可借 ggml 内核在 Mac 上实现快速本地推理,博客见 https://huggingface.co/blog/transformers-llama-cpp-quants,内核见 https://huggingface.co/ggml-org/kernels。
正文
Run GGUF models directly with transformers. This work brings ggml's Metal kernels to the transformers ecosystem, increasing compatibility and performance. More info below
Millions of GGUF downloads later, those same llama.cpp checkpoints can now run in 🤗 transformers. Same models, more ways to use them, and fast local inference on Mac powered by ggml kernels! Blog: https://huggingface.co/blog/transformers-llama-cpp-quants ggml kernels: https://huggingface.co/ggml-org/kernels在 X 查看被引用的帖子
来源:Georgi Gerganov · x.com