跳到正文

#视频

今日 14 条
9月28日周一
9月27日周日
  1. Hugging Face Daily Papers42

    DataMagic:通过声明式多智能体编排创作数据视频

    DataMagic 提出一种从原始表格数据自动生成数据视频的声明式多智能体编排方法。其 DVSpec 规范统一图表、旁白与动画并保证数据溯源,配合“先生成后编排”策略优化叙事连贯性。在 109 个真实样本上,DataMagic 将质量分数从 GPT-5 的 2.13/5 提升至 3.89(+83%),执行成功率超 95%,并将创作耗时降低 79.7%。

9月26日周六
9月25日周五
9月24日周四
  1. MiniMax (official)29

    用 MiniMax-H3 构建的更多方式。🐮 很高兴看到模型优化与 AMD 推理工程结合,为创作者带来更快的反馈循环。⚡️

    引用Nunchux AI@NunchuxAI

    5s of video, generated in 1.3s! Nunchux brings @MiniMax’s MiniMax-H3 to @AMD MI355X with up to 26.7× faster inference than SGLang on 8 GPUs @AMDServer And with streaming generation, you can change the prompt as the video plays, steering what happens next. One stack, optimized across hardware. Free access to MiniMax-H3 is coming soon! Join the waitlist now at http://nunchux.ai. Our blog: https://www.nunchux.ai/blog/video-generation-on-amd-mi355x #AI #VideoGeneration #MiniMaxH3

9月23日周三
  1. Josh Woodward42

    重大里程碑! @FlowbyGoogle:每月有超过 2500 万人使用 Google Flow 来构思新点子、创作故事、打造酷炫作品。 谢谢大家。 我们将继续为所有用户提供每天额外 50 个积分。快来继续创作吧。

    引用Google Flow@FlowbyGoogle

    More than 25 million people are using Google Flow every month to dream up new ideas, create stories, and build cool things. Thank you. We’re continuing the 50 additional daily credits for all users. Dive in and keep creating.

9月21日周一
9月20日周日
9月19日周六
9月18日周五
  1. MiniMax (official)34

    Nunchux AI 与 MIT、CMU、UC Berkeley、斯坦福及 NVIDIA 研究者合作推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上比 FlashAttention-4 快 1.6×、B300 上快 1.5×,保真度优于 SageAttention2。

    引用Nunchux AI@NunchuxAI

    Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.

9月17日周四
9月16日周三
  1. Runway News40

    Runway 详解实时视频生成的同步审核系统

    Runway 为即将发布的实时视频生成模型设计了同步审核系统,采用 Zentropi 的 CoPE-B 模型,安全审查平均耗时不到 0.5 秒,可将违规内容暴露窗口压缩至半秒以内。CoPE-B 是基于 Gemma-4-26B-A4B-it 的 LoRA 适配器,该 MoE 模型总参数 25.2B、每次前向仅激活 3.8B,兼顾速度与准确率。实时模型输出将附带 C2PA 溯源信号。

9月15日周二
9月12日周六
  1. ViggleAI42

    你肯定没见过这个: GPT-6 Astra 用于场景与道具建模 + PINOC mcp 用于可动画的高斯泼溅角色

    引用PINOC@Viggle_PINOC

    GPT-6 Astra can now generate animatable Gaussian Splat characters. We connected it to the PINOC MCP and asked for a backrooms-style, Exit 8-ish game. We described the character we wanted and the motions. [MCP link in the comment 👇] Astra generated the character and every motion through PINOC through free preset animations and text to animation, and wrote the loop and the anomaly logic itself, and shipped the whole thing in a few sessions.

9月11日周五
  1. Runway News53

    Runway 详解实时视频生成研究:从 Gen-4.5 后训练到流式输出

    Runway 分享其实时视频生成研究的整体思路,目标是把视频模型从整段生成转向优化首帧时间和随提示流式输出。方法上基于 Gen-4.5 等基础模型做后训练,先将架构改为时间因果、逐帧自回归生成,再通过两阶段分布匹配蒸馏(先 off-policy 后 on-policy)把每帧压缩到几步推理,配合渐增序列长度课程减少误差累积,同时针对并发会话构建推理服务栈。

9月9日周三
9月5日周六
  1. Runway News48

    Runway 推出 Team 团队协作套餐

    Runway 发布 Team 套餐,面向协作创作团队,提供最多 100 个共享项目、Brand Kits 和 1TB 存储,每个席位每月向统一共享余额贡献 6,900 credits,未用完的 credits 可结转一个月。该套餐今日上线网页端,每席位每月 69 美元,按年付费为 55 美元(8 折),支持 2 至 9 人团队;Standard、Pro 和 Max 转为个人套餐。

9月2日周三
  1. Runway News35

    Miro 如何用 Runway 为四地市场制作 Canvas26 主题演讲视频

    Miro 品牌团队完全用 Runway 生成 Canvas26 年度活动开场视频,并针对 4 个城市分别制作本地化版本,所有城市景观、人群和入场镜头均由 Runway 生成,再叠加动效与品牌元素。团队用场馆照片作为 Runway 参考图生成对应地点的电影感入场镜头,并借助 Runway 原生生成非标准宽高比画面,避免 4K 素材裁切损失分辨率。

  2. Fei-Fei Li57

    Fei-Fei Li 宣布 World Labs 团队发布从零训练的多模态世界模型 Atlas。该模型支持像素级精准的相机控制生成帧、从单张输入图像重建大场景、通过重排视频帧模拟时空、从一或多张图像原生输出 3D 空间,以及将多张带位姿图像组合成一致的 3D 世界,潜在应用覆盖 VFX 到机器人。

    引用World Labs@theworldlabs

    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

  3. Google DeepMind68

    Google DeepMind 为 Gemini 推出 agentic video understanding 视频分析功能

    Google DeepMind 在 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 上推出 agentic video understanding,动态检索视频片段以替代固定帧率处理,token 消耗最多降 88%,成本最多降 66%,准确率最多提升 7%。

    推荐理由:官方给出 token、成本与准确率的具体降幅和开启方式,长视频处理场景下可评估是否切换处理模式。

8月28日周五
8月25日周二
  1. OpenRouter Announcements60

    OpenRouter 发布 Video Generation API 代码指南,一个端点调用 Seedance、Veo、Wan

    OpenRouter 发布视频生成 API 教程,通过 POST /api/v1/videos 提交任务、轮询状态、下载 MP4,同一集成可用 Seedance 2.0、Veo 3.1、Wan 2.7,切换模型只需改 model 标识符。

    推荐理由:OpenRouter 官方教程给出完整的异步视频生成集成代码,包括轮询、失败状态处理和换模型时需注意的参数差异,可直接迁移到自己的项目。

8月21日周五