跳到正文

#视频

今日 13 条
今天10月2日周五
  1. Hugging Face Daily Papers43

    视频大模型时序推理为何在输出层“褪色”:TAI 方法无需训练即可增强时序表征

    视频大语言模型(VideoLLMs)的时序推理能力在中间层达到峰值,却随层数加深逐渐衰减至输出层,导致反转帧序后预测结果往往不变。研究者据此提出 Temporal Activation Injection(TAI),在峰值层提取时序表征并注入后续层,无需训练即可在三个 VideoLLM 和四个基准上稳定提升时序推理,且对非时序任务影响极小。该研究已被 NeurIPS 2026 接收。

  2. AYi65

    作者引用 Ben Affleck 在闭门峰会的分享,其创立的 16 人后期 AI 工作室 InterPositive 于 2026 年 3 月被 Netflix 以 5.87 亿美元现金全资收购。做法是解冻开源视频权重、专训最后的电影层,并用在受控舞台实拍 8 个月的私有数据集做后期训练;每部新片在自己的拍摄素材上微调专属私有模型,素材与模型迭代成果留在剧组手里。

    引用Rohan Paul@rohanpaul_ai

    Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)

    推荐理由:原文梳理了 Ben Affleck 用私有实拍数据微调开源视频模型的思路与产权闭环,读者可以借此对比公共模型与影视级生产的差距。

  3. Rohan Paul64

    Rohan Paul 对比 Ben Affleck 的前后反差:Affleck 在 2026 年 2 月称 AI 只是类似 VFX 的工具、写不出有意义的东西,随后却打造了 VFX 级 AI 工具并以 5.87 亿美元售出。作者引用的上下文称 Affleck 于 2022 年创立电影后期 AI 公司 InterPositive,通过解冻权重微调开源视频模型、仅训练最后一层电影级参数,并用自摄 8 个月数据做后期训练,Netflix 于 2026 年 3 月以 5.87 亿美元现金收购该公司。

    引用Rohan Paul@rohanpaul_ai

    Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)

  4. Thomas Wolf53

    Thomas Wolf 发推调侃 Karpathy 从 X 消失后,Ben Affleck 开始讲微调方法,称要先冻结基座权重、学习率用 2e-4。引用内容介绍 Affleck 通过解冻权重、只训练最后的电影级层来微调开放视频模型,其创办的 InterPositive 自建 8 个月数据集,并被 Netflix 以 5.87 亿美元现金收购。

    引用Rohan Paul@rohanpaul_ai

    Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)

  5. Hugging Face Daily Papers36

    4Director:用刚性 3D 几何控制视频世界模型

    4Director 是一个以显式 4D 场景表示为条件的视频世界模型,将每个物体从输入图像重建为规范网格,并按每帧一个刚性变换驱动其运动。它把受控场景渲染为深度视频,再通过 Motion Adapter 合成视角一致的外观、光照与非刚性动态。团队构建了含 20,774 个片段的 RealCOD-Rigid 数据集,并提出 IG-IoU 指标,实验显示其在视觉质量与相机、物体控制上均优于此前方法。

  6. IT Home19

    苹果 Final Cut Pro 13 曝光:支持 Pixelmator Pro 往返编辑、多语言转录等

    苹果 Final Cut Pro 13 部分新功能提前泄露,有望加入多语言转录、Pixelmator Pro 图像往返编辑、自然运动模糊、捏合缩放、剪辑断开和关键帧缓动等功能。其中往返编辑可将用户在 Pixelmator Pro 中修改的画面重新插入视频项目,不再局限于静态图像编辑。相关视频由 YouTube 用户 The Final Cut Bro 上传后已设为私有。

  7. Emad54

    Tavus 发布 Griffin,称其为首个通过视频图灵测试的模型,48% 与其实时对话的人以为是真人,此前系统通过率低于 3%,并在 NVIDIA 全双工 AI 视频基准上排名第一。Tavus 称其为首个 Human Interaction Model(HIM)。转发者 Emad Mostaque 评论称,能在屏幕另一端完成的工作 AI 能做得更好。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

10月1日周四
  1. 赵纯想38

    赵纯想指出,Opus 5.5 生成的精美视频第一天被欢呼、第二天被复刻、第三天便索然无味,原因是 Skill 的分发把差异拉平,而差异正是“力”与利润率的来源。他引用红果短剧数据称,95% 以上的漫剧亏本,头部若干漫剧利润率百分之几万,赚走大部分钱。AI 把个体生产力拉到三年前不敢想象的水平,但盈利仍遵循极致冷酷的大数定律。

    引用赵纯想@chunxiangai

    力,起源于差异。一大一小,两个同样材质的东西,能在人脑中映射出“压迫感”。这个压迫感,就是力。把房子建在房子群落中,房子就不够有力。把房子建在旷野上,方圆百里,没有其他房子,这个房子就有力。力,就是差异。 有了力,就有了结构。一大一小,就是一种结构。荒野上的孤独的房子,就是结构。结构就是稳态的差异。海浪也有差异,但不稳定。海浪,不是结构。结构若想稳定存在,就要形成系统。系统保存着结构,结构保存着差异,差异,保存着力。 少部分人,发抖音。大部分人,刷抖音。 一少一多,就有力。 少部分人,购买奢侈品。奢侈品每年花大价钱,买最大的广告牌,让大部分人看到、得知它的贵重。 一多一少,就有力。 越有力的系统,结构上的差异越大。女人,细枝挂硕果,最有力。男人,毛发旺盛,但粉鸡巴,最有力。 结构上的差异越大,包含这个结构的系统,就越难以自持。抖音,通过算法控制,来维系力。让少部分精彩视频,被分发给大部分人。一多一少,让力释放其价值。没有推荐算法,任由内容传播,会使得精彩和平庸混为一谈,那是对力的亵渎。 越难自持的系统,就需要越强的外部能量的控制。在强力的控制下,围绕力,释放价值,价值反哺外部能量,实现自持。如同涡扇发动机,进入自持,需要外部能量的点火激励。 用草纸,画出一个二八定律下的结构,一个有力的平台,便被描绘出来。少部分人,做什么。多部分人,做什么。一多一少,一薄一厚,一尖一盾,就形成力。有力,就有价值。结构,保存力,等同于保存价值。系统,组织结构,让价值可视、可流转、可维护。 大而全,是亵渎力。小而美,是小的力。大而精,最有力。力在悬殊中,力在差异里。

  2. MiniMax (official)55

    HeyGen 发布 HeyGen Video,面向需要制作级视频但不想承担制作级成本的企业,基于 MiniMax H3 构建并由 HeyGen 后训练。10 月前定价 $0.01/s(五折优惠),详情见 https://developers.heygen.com/heygen-video-1.0-catalog。MiniMax 官方转发并表示自豪提供基础模型。

    引用HeyGen@HeyGen

    We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog

  3. MiniMax (official)34

    基于 MiniMax H3,@Creatify_Labs 的 Boreal-H3 是一款专为广告优化的视频模型,在保持产品和角色一致性的同时,更准确地遵循创意简报。 期待看到 MiniMax H3 成为更多面向特定行业的前沿模型的基础!✨

    引用Creatify Labs@Creatify_Labs

    Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.

9月30日周三
  1. Kling AI40

    满血版 KLING 4.0 实战演示 🎬 #Kling4 #KlingAI #KlingModeOn

    引用汗青 HQ@hq4ai

    SPARE。由 可灵 Kling 4.0 满血版生成的短片。 感谢可灵 AI 的内测邀约。本片使用 可灵 Kling 4.0 满血版全能参考生成。 满血版比 Flash 还是强不少。 - 原生 30 秒直出 - 画面与声音更清晰 - 口型匹配更精准 - 21:9 电影画幅 可灵 Kling 4.0 将于 10 月上线。 中文字幕版:

  2. Hugging Face Daily Papers39

    LEAP:面向长音视频感知的学习式分块证据检索框架

    LEAP 通过将录音切分为固定时长块、用轻量定位阶段为每块候选窗口打分,再把高分窗口汇聚后单次有界重编码作答,使答案输入与峰值上下文不随录音时长增长。在多个 AVQA 基准上,LEAP 较 Qwen3-Omni-30B-A3B 基线提升 4.5-16.8%,迁移到 MiniCPM-o 4.5 后超出其已发表结果 3.1-13.0%。

  3. Hugging Face Daily Papers34

    MemLife:面向长期第一视角视频记忆的整理与推理系统

    MemLife 是一个多模态记忆系统,通过构建实体锚定的第一人称文本片段,并用时间索引的智能体读取器进行检索,在无需训练、查询时不访问原始视频的情况下,于四个长时程基准上比最强免训练基线提升 4.6–12.0%。配套的 MemOpt 强化学习框架优化记忆写入器,使 MemLife 再提升 2.7–5.0%,且增益可跨写入器与读取器骨干及不同记忆系统泛化。

  4. Hugging Face Daily Papers40

    StreamMAE:让自监督学习在连续视频流上奏效

    研究团队构建了 95 小时城市步行游览视频数据集 WT++,用于严格按时间顺序、滑动窗口批次的流式自监督预训练。结果显示对比学习和蒸馏方法在此设定下表现不佳,MAE 较稳健但仍不及标准 i.i.d. 预训练,主要瓶颈是批内帧高度相似。

9月29日周二
9月28日周一