跳到正文
@AntLingAGI· @AntLingAGI · X·· 2026-09-05AI 评分42
AI 导读

视频理解与多模态智能体 Ling-3.0-flash-VL 会拆解请求,在视频时间线上搜索,对候选片段进行推理,并找到目标片段。随后它可以调用工具提取关键帧、裁剪图像并验证结果。

正文

Video Understanding and Multimodal Agents

Ling-3.0-flash-VL decomposes a request, searches across the video timeline, reasons over candidate clips, and finds the target segment. It can then call tools to extract keyframes, crop images, and verify the result.

来源:@AntLingAGI · x.com