跳到正文

#多模态

今日 25 条
9月30日周三
9月29日周二
  1. Microsoft Research61

    Microsoft Research 发布生物研究领域 AI 系统 Quine

    Microsoft Research 推出 Quine,一个面向生物学的多模态世界模型与交互式 harness,连接科学工具、文献和研究人员。在与 Broad Institute 合作中,Quine 用于预测可驱动胰腺癌肿瘤细胞状态转变的化合物,排名第一的候选化合物在湿实验中产生了最大的预期细胞状态转变,从缩小化合物范围到确定候选名单仅用了一个周末。

    推荐理由:官方披露了系统构成和胰腺癌湿实验验证结果,读者可以据此评估AI世界模型在生物研究中的实际作用。

  2. OpenBMB30

    MiniCPM-o 4.5 现已支持 SGLang Omni v0.1.7。 为开发者提供更多灵活的运行和构建方式。

    引用Guitar Cat + LLM@GenAI_is_real

    Hi everyone, today we released SGLang Omni v0.1.7. This release includes 75 merged PRs and welcomes 8 new contributors, with 8 first-time contributions. We added MiniCPM-o 4.5, NVIDIA PersonaPlex-7B, and OmniTyper powered by MLX streaming ASR, while further improving realtime and stateful Omni serving. 1.Performance: continued optimizations for Qwen3-TTS, Qwen3-Omni, CosyVoice3, MOSS-TTS, and AuK, covering Prefill CUDA Graph, speaker/reference encoding, kernel fusion, batching, and vocoder hot paths. 2.Serving: added Omni session lifecycle, the SGLang streaming session bridge, and a shared /v1/realtime WebSocket runtime, while further improving realtime ASR and streaming serving. 3.Models & hardware: added MiniCPM-o 4.5 multimodal input and speech output, plus PersonaPlex-7B offline speech-to-speech. MiniCPM-o and MiniMax-Music3 now support Intel XPU, with further MUSA support for Qwen3-TTS. 4.Runtime: improved breakable Prefill CUDA Graph, Talker / Code2Wav colocation, priority CUDA streams, scheduler admission, and profiling infrastructure to reduce host overhead and improve high-concurrency stability. https://github.com/sgl-project/sglang-omni/releases/tag/v0.1.7 https://github.com/sgl-project/sglang-omni

  3. clem 🤗77

    AMD 宣布欢迎 World Labs 和李飞飞加入 AMD,双方计划结合 World Labs 在 AI 与世界模型方面的专长与 AMD 的算力能力,推动 AI 未来并强化开放 AI 生态。Hugging Face CEO Clément Delangue 转发该消息并祝贺,期待双方未来数年的成果。

    引用Lisa Su@LisaSu

    So excited to welcome @theworldlabs and @drfeifei to the @AMD family! I’ve always been a huge fan of Fei-Fei and her pioneering research in AI. Together, we’ll combine World Labs’ deep expertise in AI and world models with AMD’s compute leadership to power the future of AI and strengthen the open AI ecosystem. Can’t wait for all we’ll accomplish!

    推荐理由:AMD 收购 World Labs 与李飞飞的消息结合 World Labs 专注世界模型与开放生态的定位,读者可了解这次结合对 AI 开源生态的影响。

  4. Hugging Face Daily Papers35

    音频-视觉大语言模型中的“问题接力”机制与 SECRET 去幻觉方法

    音频-视觉大语言模型(AVLLM)存在源混淆接地幻觉,即未使用模态的线索会诱发所需模态不支持的回答。路径干预与表征分析揭示“问题接力”机制是成因,据此提出的免训练方法 SECRET 通过对比不同模态路径干预诱发的问题表征,将问题状态引向所需源证据。

  5. Hugging Face Daily Papers40

    LoopVL:循环视觉智能

    LoopVL 将 Loop Transformer 扩展到视觉语言模型,通过 Module-Loop 与 Model-Loop 计算,用共享模块迭代更新统一的视觉语言状态。该模型从零开始经语言预训练、多模态训练和后训练,在多模态理解与视觉推理基准上超过同规模及更大规模的非循环模型。研究还观察到 LoopVL 出现"视觉顿悟时刻",即视觉注意力随循环次数发生显著转移。

9月28日周一
  1. elsewhere articles41

    群核科技用空间智能重建永泰龟城,24 亿高斯点刷新全球 3D 高斯重建纪录

    群核科技联合博主「特能斯」4 人 4 天在甘肃永泰龟城采集 6 万多张照片,通过 3D 高斯重建平台 Aholo Reality 生成 24 亿高斯点、60 万平方米的数字古城,刷新全球公开可查的 3D 高斯重建纪录。该数字场景已交给景泰当地文旅作为永久数字文化档案保存,并可通过 Aholo Reality 平台在线漫游;同一组 3D 场景还被用于 LuxReal 生成 AI 短剧《永泰无战事》。

9月26日周六
9月25日周五
  1. karminski-牙医59

    美团 LongCat 发布 LongCat-2.5-Preview 模型,总参数 1.6T、激活约 48B,支持 1M token 上下文窗口,原生多模态,面向终端、浏览器、GUI、表格和设计工具等长程任务。API 与聊天入口分别见 https://longcat.ai/platform/ 和 https://longcat.ai/chat/;作者补充定价与之前一样,图中显示输入(缓存未命中)2.00 元/百万 token,输入(缓存命中)0.04 元/百万 token,输出 8.00 元/百万 token。

    引用Meituan LongCat@Meituan_LongCat

    LongCat-2.5-Preview is now live. 1.6T parameters. ~48B active. A 1M-token context window. Natively multimodal. Built to take on long-horizon tasks. From terminals and browsers to GUIs, spreadsheets, and design tools. Try it now: 🚀 API: https://longcat.ai/platform/ 💬 Chat: https://longcat.ai/chat/

  2. Google DeepMind51

    Google DeepMind 推出 Gemini 3.8 Live with Live Avatar

    Google DeepMind 推出 Gemini 3.8 Live with Live Avatar,为对话模型加入近实时视频生成与视觉形象,今日起在 Gemini Enterprise 提供。该功能支持精确唇形同步、异步工具调用不中断对话、跨 97 种语言的语音到语音同步,并可用参考图定制头像(需企业允许名单),所有输出带 SynthID 水印。

9月24日周四
  1. Karina31

    当 @jakubzeg 和我在 OpenAI 相遇时,我们反复回到一个共同的困扰:太多 AI 产品都从一个空白的聊天框开始。你几乎可以做任何事,但你必须决定从哪里开始,这让人不知所措。 有了 ACTx486,我们从一段本身已有故事的媒体内容出发,让你在观看的同时与之互动。这也让这项技术更加通用。 我们在这里写了更多关于这一选择的思考:https://www.actx486.com/

    引用Rehan Sheikh@rehan_shei

    interactive media will be huge in the next year or two! this is incredible

9月23日周三
  1. Google DeepMind68

    Google DeepMind 发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 语音生成模型

    Google DeepMind 发布 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 两款文本转语音模型,支持用自然语言提示词设计角色声音、30 秒音频样本复刻声音,并逐行指挥表演。

    推荐理由:原文给出两个 TTS 模型的定位差异、基准成绩和开放渠道,读者可据此评估语音生成工作流的选型。