AI 导读
Perplexity 开源本地推理引擎 Lily,用于在 Apple Silicon 上本地运行模型,并支撑 Mac 应用新推出的混合计算功能。该引擎针对 Qwen3.6-35B-A3B 做了专门优化,目标是让端侧计算不成为 Perplexity Computer 任务的瓶颈。图中数据显示,在 M5 Max 40 核 GPU 与 128 GB 统一内存、Q4 量化、batch 1 条件下,其 prefill 吞吐为 4156 tokens/s、decode 吞吐为 170.0 tokens/s,分别为 MLX-LM 的 1.23 倍和 1.35 倍。
正文
We're open-sourcing Lily, Perplexity's local inference engine for serving models locally on Apple Silicon. This powers Perplexity's newly introduced hybrid compute feature for the Mac app. https://t.co/RPnwLOHsSk https://t.co/yV7ryTsLqI
Today we’re open-sourcing Lily, the local inference engine we built for hybrid compute in Perplexity Computer. Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, built so on-device compute doesn’t bottleneck Computer tasks. Read more: https://t.co/OnowTP3ql6在 X 查看被引用的帖子
来源:@AravSrinivas · x.com