AI 导读
视频理解与多模态智能体 Ling-3.0-flash-VL 会拆解请求,在视频时间线上搜索,对候选片段进行推理,并找到目标片段。随后它可以调用工具提取关键帧、裁剪图像并验证结果。
正文
Video Understanding and Multimodal Agents
Ling-3.0-flash-VL decomposes a request, searches across the video timeline, reasons over candidate clips, and finds the target segment. It can then call tools to extract keyframes, crop images, and verify the result.
来源:@AntLingAGI · x.com