阿里巴巴发布一体化视频生成与编辑模型 Wan 3.0,可生成最长 30 秒 1080p 带原生音频的视频,并接受文本、图像、视频、音频、文档和网页作为创作参考。
第三方竞技场盲测排名给出 Wan 3.0 与同类视频模型的直接对比,并列出各分辨率定价。
Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audio
Wan 3.0 is Alibaba's new all-in-one video generation and editing model, positioned as a single system for turning multimodal creative direction into video. It generates up to 30 seconds at 1080p with native audio and accepts text, images, video, audio, documents, and web pages as creative references. The same model supports Text to Video, Image to Video, reference-based generation, and instruction-led editing, including changes to visuals, plot, dialogue, and sound.
In the Artificial Analysis Video Arena, Wan 3.0 ranks #1 in Video Editing with Audio, #2 in Text to Video with Audio, and #5 in Image to Video with Audio.
Wan 3.0 marks a large generational improvement: against the most recent Wan 2.7 version on each leaderboard, it rises from #5 to #1 in Video Editing with Audio, #6 to #2 in Text to Video with Audio, and #12 to #5 in Image to Video with Audio.
Wan 3.0 is available now in public preview through Alibaba Cloud Model Studio. Pricing starts at $0.05 per second for 480p, increasing to $0.10 for 720p and $0.20 for 1080p.
Congratulations to @Alibaba_Wan and @alibaba_cloud on the release!
See below for comparisons between Wan 3.0 and other leading models in the Artificial Analysis Video Arena 🧵
来源:@ArtificialAnlys · x.com