跳到正文
@AIatMeta· @AIatMeta · X·· 13 天前AI 评分39
AI 导读

实时视频流需要即时响应,同时在长时间对话中保持视觉一致性。 我们通过将一个带有 3-way CFG 的大型 40 步扩散模型教师(每个视频块 120 次评估)蒸馏为一个无引导的 2 步因果学生模型,并配备固定长度 KV cache,实现了这一目标。 Self-forcing 帮助学生模型抵抗漂移,以 60 倍更少的评估次数维持接近教师模型的质量。

正文

Live video streaming needs to respond instantly while remaining visually consistent over long conversations.

We achieved this by distilling a large 40-step diffusion teacher with 3-way CFG (120 evaluations per video chunk) into an unguided 2-step causal student with a fixed-length KV cache.

Self-forcing helps the student resist drift and maintain near teacher quality with 60x fewer evaluations.

来源:@AIatMeta · x.com