Google DeepMind 发布 Gemini Omni,称这是迈向能从任意输入创建任意内容模型的第一步,先从视频生成开始。官方表示该模型结合 Gemini 的智能与自身生成式媒体系统,在世界理解、多模态与编辑能力上有所推进。转发这一消息的 X 用户列出五项能力,包括符合现实的物理模拟、角色面部保持一致、用自然语言改背景换人物加特效,并称可在 Gemini 应用试用 Omni Flash。
Damn! Google has really gone absolutely wild this time. Gemini Omni is about to blow the roof off the ceiling of video generation 🤯
Making videos used to be like building with Lego blocks, piece by piece, slowly.
Now it’s giving you a magic Lego factory that can actually think.
You chat in natural language, and it understands real-world physics, history, biology, culture—then directly generates or edits any video.
Five most mind-blowing abilities that you can use right now:
1Understands real physics—glass marbles colliding, turning, and bouncing in ways that match reality.
2Faces never get distorted—define a character once, put them in any scene, any action.
3Edit videos like you edit ChatGPT text—change backgrounds, swap people, add effects with a single sentence.
4Upload an image and apply any style—make claymation, visualize protein folding, whatever you imagine.
5Video isn’t a dead file anymore—change angles, lighting, objects, even storylines just by chatting.
This isn’t a competitor to Sora.
This is the first time a world model has truly entered a consumer-facing product.
It’s not just generating pixels—it’s simulating a coherent physical and semantic world.
Open the Gemini app right now and try Omni Flash.
Go try it. You’ll thank me later.
Video
We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 Video在 X 查看被引用的帖子
来源:@AYi_AInotes · x.com