🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE
X:谢赛宁
@sainingxie · X
切换来源
Saining Xie@sainingxieAI 评分4545引用Minghui Guo@MinghuiGuo77
Saining Xie@sainingxieAI 评分2424引用David@DavidSHolz@scaling01 i actually filmed and edited the video myself 😅 also implemented the 3d visualization technical deep dive, it all runs in a browser in realtime!
Saining Xie@sainingxieAI 评分4141大脑如何从(可能不完整且有噪声的)视觉观察中构建并追踪世界的内部状态? 我相信视觉状态追踪将成为未来几年视觉领域的重大挑战,希望这个基准能成为一个有用的起点。enjoy!
引用Sihyun Yu@sihyun_yuCan MLLMs actually track what's happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. https://vision-x-nyu.github.io/vstat-site/ 🧵 [1/11]
Saining Xie@sainingxieAI 评分3737引用Jihan Yang@jihanyang13Camera pose matters for video understanding! Today's MLLMs excel at recognizing activities, but still struggle with the underlying space and ego/object dynamics in video. We trace this gap to a missing piece: camera pose. Introducing Cambrian-P: a multimodal LLM natively grounded in camera pose. (1/n)
Saining Xie@sainingxieAI 评分4343引用Jaskirat Singh@1jaskiratsinghIn Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation. Today we make them a bit more simple. a bit more general. Result: >10x faster convergence, better reconstruction, better generation. And yes we test them on T2I and world models :) Introducing RAEv2
Saining Xie@sainingxieAI 评分1414说到最爱的滤波器…… 最初的出现,大多是偶然 @TacoCohen: 我记得那些日子,一次成功的实验就是能产出酷炫的滤波器

引用Taco Cohen@TacoCohenI remember the days when a successful experiment was one that produced cool filters