跳到正文
@Alibaba_Qwen· @Alibaba_Qwen · X·· 19 天前精选AI 评分69
AI 导读

通义千问发布 Qwen3.8-Omni-Flash,这是其首个围绕智能体能力构建的全模态模型,把音视频理解、推理与工具调用整合在同一模型中。官方称其音视频能力接近 Gemini 3.8 Flash,在 WildClawBench-MM 与 UniClawBench 上的智能体表现平均提升 19.5 分。

推荐理由

官方列出了与 Gemini 3.8 Flash 的能力对比和 token、成本降幅,读者可据此判断长音视频智能体工作流的可用性。

正文

🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities!

Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result.

Highlights: 🥳
- Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps.
- A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench.
- 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench.

Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever.

To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️

We can't wait to see what you build with Qwen3.8-Omni-Flash! 👀

- Blog: https://t.co/oM9V1TkqYF
- Qwencloud: https://t.co/Cc2I8ELAnD
- Qwen Studio: https://t.co/V7RmqMaVNZ
- API: https://t.co/lNE7fH5YUt
- Qwen-MM-Plugins: https://t.co/SnM27dDP3d
- Qwen-Live Harness: coming soon
https://t.co/iVYlGjIbdy

来源:@Alibaba_Qwen · x.com