AI 导读
智能体每一轮都会重读自己的全部历史。复用模型对这段历史的缓存,就能跳过这项工作。把这份缓存存到 GPU 之外原本很浪费:每个智能体各自保存一份副本,而且每一轮都从头重新保存全部内容。现在,每个共享前缀只保存一次,每一轮只写入新增的部分。(2/5)
正文
Agents reread their whole history every turn. Reuse the model's cache of it and you skip that work. Saving that cache off-GPU was wasteful: agents each saved their own copy, and every turn re-saved everything from scratch. Now, one save per shared prefix, and each turn writes only what's new. (2/5)
来源:@SemiAnalysis_ · x.com