跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 14 天前AI 评分43
AI 导读

谷歌新论文指出,在无服务器 CPU 上运行小型量化 LLM 时,55–70% 的冷启动延迟仅仅来自模型加载。 也就是说,瓶颈往往在于搬运模型权重,而非生成 token。 推理本身的问题,还不如把模型载入内存来得大。 给同一个模型分配 8 GB 的 Cloud Run 内存而非 4 GB,可解锁约 2 倍 CPU,将热推理时间几乎减半。

正文

New Google paper shows for small quantized LLMs on serverless CPUs, 55–70% of cold-start latency is just model loading.
i.e. the bottleneck is often moving model weights, not generating tokens.

The inference itself is less of the problem than getting the model into memory.

Giving the same model 8 GB of Cloud Run memory instead of 4 GB unlocks roughly 2× the CPU, cutting warm inference time almost in half.

来源:@rohanpaul_ai · x.com