AI 导读
三种模式支持实际部署。 daemon 加载并持有缓存 client 连接到已有的 daemon off 使用标准磁盘路径 如果 daemon 崩溃,已映射的引擎继续运行。CUDA 只有在所有引用都消失后才释放内存。
正文
Three modes support practical deployment.
daemon loads and owns the cache
client connects to an existing daemon
off uses the standard disk path
If a daemon crashes, mapped engines keep running. CUDA frees the memory only after every reference is gone.
来源:@AntLingAGI · x.com