VTR-Bench:视频生成中视觉文本渲染能力的系统评测基准
VTR-Bench 提出一套系统评测基准,专门评估视频生成模型的视觉文本渲染能力,包含 300 个覆盖广告、科学视频等五类场景的提示词,并开发了与人工对齐的自动化评测流程。
VTR-Bench 提出一套系统评测基准,专门评估视频生成模型的视觉文本渲染能力,包含 300 个覆盖广告、科学视频等五类场景的提示词,并开发了与人工对齐的自动化评测流程。
OneStreamer提出通过共享的主动生成过程,联合学习查询无关的证据记录与任务响应,其主动分层字幕记忆(PHCM)生成带时间锚定的局部细节字幕与已完成事件摘要,在不回溯历史视觉特征的情况下提供可复用的事实上下文。
视频大语言模型(VideoLLMs)的时序推理能力在中间层达到峰值,却随层数加深逐渐衰减至输出层,导致反转帧序后预测结果往往不变。研究者据此提出 Temporal Activation Injection(TAI),在峰值层提取时序表征并注入后续层,无需训练即可在三个 VideoLLM 和四个基准上稳定提升时序推理,且对非时序任务影响极小。该研究已被 NeurIPS 2026 接收。
Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)
推荐理由:原文梳理了 Ben Affleck 用私有实拍数据微调开源视频模型的思路与产权闭环,读者可以借此对比公共模型与影视级生产的差距。
Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)
Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)
World Observer 通过联合生成一个 actor 视角与一个或多个全景 observer,让离开 actor 视野的物体在 observer 中持续演化,从而在重新进入视野时保留其状态与动态。
PROWBench 是一个评估视频模型对程序指定世界事件还原保真度的基准,包含 170 个程序化构建的片段和 600 段代理视频,记录实体状态与带时间戳事件(含镜头视野外事件)作为可重放世界记录。
4Director 是一个以显式 4D 场景表示为条件的视频世界模型,将每个物体从输入图像重建为规范网格,并按每帧一个刚性变换驱动其运动。它把受控场景渲染为深度视频,再通过 Motion Adapter 合成视角一致的外观、光照与非刚性动态。团队构建了含 20,774 个片段的 RealCOD-Rigid 数据集,并提出 IG-IoU 指标,实验显示其在视觉质量与相机、物体控制上均优于此前方法。
苹果 Final Cut Pro 13 部分新功能提前泄露,有望加入多语言转录、Pixelmator Pro 图像往返编辑、自然运动模糊、捏合缩放、剪辑断开和关键帧缓动等功能。其中往返编辑可将用户在 Pixelmator Pro 中修改的画面重新插入视频项目,不再局限于静态图像编辑。相关视频由 YouTube 用户 The Final Cut Bro 上传后已设为私有。
Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).
Tavus 发布 Griffin,公司称其为首个面向实时面对面对话的 Human Interaction Model,可处理语音、表情、语调、手势和停顿并接收与生成视频。
Computer 里的 Seedance 对营销和品牌团队来说相当惊艳!
Generate a commercial or brand video for your small business with Seedance 2.5 in Perplexity Computer. This commercial was created entirely inside Computer, including the brand look, product mockups, and full video edit.
力,起源于差异。一大一小,两个同样材质的东西,能在人脑中映射出“压迫感”。这个压迫感,就是力。把房子建在房子群落中,房子就不够有力。把房子建在旷野上,方圆百里,没有其他房子,这个房子就有力。力,就是差异。 有了力,就有了结构。一大一小,就是一种结构。荒野上的孤独的房子,就是结构。结构就是稳态的差异。海浪也有差异,但不稳定。海浪,不是结构。结构若想稳定存在,就要形成系统。系统保存着结构,结构保存着差异,差异,保存着力。 少部分人,发抖音。大部分人,刷抖音。 一少一多,就有力。 少部分人,购买奢侈品。奢侈品每年花大价钱,买最大的广告牌,让大部分人看到、得知它的贵重。 一多一少,就有力。 越有力的系统,结构上的差异越大。女人,细枝挂硕果,最有力。男人,毛发旺盛,但粉鸡巴,最有力。 结构上的差异越大,包含这个结构的系统,就越难以自持。抖音,通过算法控制,来维系力。让少部分精彩视频,被分发给大部分人。一多一少,让力释放其价值。没有推荐算法,任由内容传播,会使得精彩和平庸混为一谈,那是对力的亵渎。 越难自持的系统,就需要越强的外部能量的控制。在强力的控制下,围绕力,释放价值,价值反哺外部能量,实现自持。如同涡扇发动机,进入自持,需要外部能量的点火激励。 用草纸,画出一个二八定律下的结构,一个有力的平台,便被描绘出来。少部分人,做什么。多部分人,做什么。一多一少,一薄一厚,一尖一盾,就形成力。有力,就有价值。结构,保存力,等同于保存价值。系统,组织结构,让价值可视、可流转、可维护。 大而全,是亵渎力。小而美,是小的力。大而精,最有力。力在悬殊中,力在差异里。
LongLive-Plug 面向视频生成的一次性通用蒸馏 paper: https://huggingface.co/papers/2609.38154
We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog
Runway 发布新产品 Runway Ads,连接广告账户和品牌套件后可生成视频与图像广告创意,直接发布到 Meta、Google 和 TikTok,并读取平台表现数据迭代下一轮创意。
推荐理由:Runway 官方发布并给出自身投放数据,读者可以据此评估端到端生成式投放工具对营销工作流的实际改变。
Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.
作为一个个人副业项目,我一直在做一款电子游戏的原型。 作为一名前游戏创始人,制作游戏的过程极具满足感和创造力,不要把这件事外包给AI。用AI与你协作,把你的愿景变为现实。 以下是我的尝试方式:
广告做完了。版本还没完。 不同尺寸。不同市场。同一广告活动。 用 Luma 的 Ad Variants 让你的广告走得更远。
满血版 KLING 4.0 实战演示 🎬 #Kling4 #KlingAI #KlingModeOn
SPARE。由 可灵 Kling 4.0 满血版生成的短片。 感谢可灵 AI 的内测邀约。本片使用 可灵 Kling 4.0 满血版全能参考生成。 满血版比 Flash 还是强不少。 - 原生 30 秒直出 - 画面与声音更清晰 - 口型匹配更精准 - 21:9 电影画幅 可灵 Kling 4.0 将于 10 月上线。 中文字幕版:
Window Lab,用 Krea Agent 创建的应用。 只需上传你的视频,调整运动控制,即可获得你想要的效果。 在下方试试 👇
Google 分享了在 TPU 上把视频扩散模型稀疏注意力从理论稀疏转化为端到端加速的案例研究。81 帧 720p 视频的序列长度达 50K 至 400K,从 720p 升到 1440p 序列长度翻两番,全注意力在单层延迟中占比可从 55.5% 升至 88.2%。
LEAP 通过将录音切分为固定时长块、用轻量定位阶段为每块候选窗口打分,再把高分窗口汇聚后单次有界重编码作答,使答案输入与峰值上下文不随录音时长增长。在多个 AVQA 基准上,LEAP 较 Qwen3-Omni-30B-A3B 基线提升 4.5-16.8%,迁移到 MiniCPM-o 4.5 后超出其已发表结果 3.1-13.0%。
MemLife 是一个多模态记忆系统,通过构建实体锚定的第一人称文本片段,并用时间索引的智能体读取器进行检索,在无需训练、查询时不访问原始视频的情况下,于四个长时程基准上比最强免训练基线提升 4.6–12.0%。配套的 MemOpt 强化学习框架优化记忆写入器,使 MemLife 再提升 2.7–5.0%,且增益可跨写入器与读取器骨干及不同记忆系统泛化。
研究团队构建了 95 小时城市步行游览视频数据集 WT++,用于严格按时间顺序、滑动窗口批次的流式自监督预训练。结果显示对比学习和蒸馏方法在此设定下表现不佳,MAE 较稳健但仍不及标准 i.i.d. 预训练,主要瓶颈是批内帧高度相似。
作者分享花费 10+ 小时摸索出的用 Opus 5.5 做视频生成的工作流:建议搭配 Claude Code 使用,用 OpenRouter API 一个密钥调用图像。
感谢 Yihui,出品速度好快呀。 正在努力让下面这种视频,在 YouMind 里也能顺畅制作。
来了 @stark_nico99 @YouMind_AI ,Manus 风格的YouMind 宣传片。
美国:250 年致敬。 我制作了这部 5 分钟纪录片,完全由 opus 编排,混剪了美国历史上最具标志性时刻的重制版。
Seedance 2.5 草稿模式现已登陆 Krea。 用 480p 生成做实验——场景合适时再切换到 1080p。 立即体验 👇
OpenRouter 比较了自家平台四条图生视频模型线的分辨率、时长、帧控制、音频和价格,覆盖 Veo 3.1、Seedance 2.5 与 2.0 系列、Kling v3.0 和 Grok Imagine Video 1.5,价格数据核查于 2026 年 9 月 11 日。
PixelUMM 是一个无需编码器的统一多模态模型,直接在像素空间完成图像与视频的理解和生成,将图像表示为空间 patch、视频表示为时空 tubelet,通过单层线性投影接入共享多模态主干。
针对音视频联合扩散模型的多奖励强化学习,研究者提出 Adaptive Reward Routing,在 DiffusionNFT 前向过程 RL 中同时自适应调整更新位置与奖励组合。
Google 与 XPRIZE、Range Media Partners 合办的 Future Vision XPRIZE 公布大奖,独立导演 Jeff Synthesized 凭借短片 The Gifted 从全球超 2500 部作品中胜出。