Gemini 的智能体视频理解开始在 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 上通过 Google AI Studio 的 API 推送,并即将登陆 Gemini App。该能力不再扫描整个文件,而是跨视频字幕、音频和画面做推理,动态调整帧率以定位所需片段,长视频内容(从 10 分钟指南到数小时录像)的效率提升最明显。官方图表显示,在 1H-VideoQA 与 LVBench 等长视频基准上,单次查询 token 从约 30 万至 40 万降至 5 万以内,准确率同步小幅上升。
官方图表给出长视频场景 token 消耗大幅下降而准确率上升的对比,读者可据此判断长视频处理的成本空间。
Instead of scanning an entire file, Gemini reasons across the video’s transcript, audio, and frames, dynamically adjusting the frame rate to pull the exact moments needed.
The efficiency gains are most significant for long-form content, from 10-minute guides to multi-hour recordings.
Agentic video understanding is rolling out to 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via API in @GoogleAIStudio and coming soon in the @GeminiApp.
Find out more → https://t.co/bjPn5V52zr
来源:@GoogleDeepMind · x.com