跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-29AI 评分40
AI 导读

伯克利与德州大学提出新推理引擎 FreeToken,让 8GB 显存游戏本以 39.3 tokens/s 运行 35B 模型,速度快于生产环境中的 Codex。它让 GPU 缓存跟随路由器实时需求,把未命中专家从 llama.cpp 的 62% 降至 16%,其余专家按机器实际 PCIe 与内存速度计算的比例,部分拷贝到 GPU、部分留在系统内存原地运行,避免双方互相等待。

正文

An 8 GB gaming laptop can run a 35B model at 39.3 tokens a second, faster than Codex in production, if the software stops guessing which parts of the model to keep on the GPU.

That's FreeToken, New research from Berkeley and University of Texas, a new engine for running big open models on ordinary machines.

Every other local engine picks which experts sit on the GPU at load time and leaves them there, so when the model routes elsewhere, that expert gets copied over PCIe while the CPU sits idle.

FreeToken lets the GPU cache follow whatever the router just asked for, cutting missed experts from llama.cpp's 62% to 16%.

The rest it splits two ways: some copied to the GPU, the others run in place in system RAM, in a ratio computed from your machine's real PCIe and memory speeds, so neither side waits on the other.

来源:@rohanpaul_ai · x.com