跳到正文

模型发布

新模型的发布、开源与迭代:大模型厂商的旗舰更新、开源权重放出、性能与价格变化的第一时间记录。

243条精选相关主题产品更新论文研究开源生态

最新精选

第 181–200 条 · 共 243 条
5月20日周三
  1. @kimmonismus82

    Gemini Omni 发布,官方介绍这是一款可从任意输入生成内容的新模型,首批聚焦视频,已在 Gemini App、Flow 和 YouTube 上线,API 支持即将推出。转发的作者 @kimmonismus 称这是真正的惊艳时刻,认为它是迈向 AGI 的世界模型,能从任何输入生成任何内容。

    引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganK

    Introducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video

    推荐理由:官方说明了 Gemini Omni 从任意输入生成内容的能力和首批视频场景,读者可据此了解当前可用渠道与范围。

5月19日周二
  1. @berryxia66

    Odyssey 发布 Agora-1,一个多智能体世界模型,人类与 AI 可同时进入同一模拟世界并实时互动、互相影响。官方推出可游玩的研究预览,用 Agora-1 模拟多人 GoldenEye 死亡竞赛,模型实时生成画面和声音,整个世界持续更新。

    引用Odyssey (@odysseyml)@odysseyml

    Introducing Agora-1, a multi-agent world model. Multiple participants—human or AI—can now interact inside the same world simulation, all in real-time. Try our playable research preview today, with Agora-1 simulating a multiplayer GoldenEye deathmatch! Video

    推荐理由:世界模型从单人视频生成扩展到多人实时共享模拟,读者可据此了解人机共处同一模拟世界的当前形态。

5月18日周一
5月16日周六
5月15日周五
  1. @vista866

    面壁智能发布 1.3B 参数的视觉模型 MiniCPM-V 4.6,面向消费级和移动硬件,已在 Hugging Face、GitHub 和 ModelScope 上线。该模型采用 LLaVA-UHD v4 技术,官方称将视觉编码成本降低 55%。官方还称其在多模态和 Artificial Analysis 基准上超过 Gemma4-E2B-it 和 Qwen3.5-0.8B,TTFT 为 75.7ms、比 Qwen3.5-0.8B 快 2.2 倍。原文作者表示在 Hugging Face 看到该模型论文,准备抽空测试。

    引用OpenBMB (@OpenBMB)@OpenBMB

    1/5 MiniCPM-V 4.6 (1.3B) is now live 🚀🚀 High-res visual processing, optimized for consumer-grade and mobile hardware. We’ve leveraged the latest LLaVA-UHD v4 technique to cut vision encoding costs by 55%, enabling native edge deployment with extreme efficiency. 🔥 Beats Gemma4-E2B-it and Qwen3.5-0.8B across key multimodal and Artificial Analysis benchmarks — scoring higher than Qwen3.5-0.8B using just 2.5% of its token budget. ⚡ TTFT (75.7ms) 2.2x Faster than Qwen3.5-0.8B even with 3136² high-res images. 🏗️ ~1.5x Token Throughput compared with Qwen3.5-0.8B on a single RTX 4090. Try the model here: 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM-V 🔭 Modelscope: modelscope.cn/models/OpenBMB… 🌐 Web Demo: huggingface.co/spaces/openbm… 📱 App Demo: github.com/OpenBMB/MiniCPM-V… Video

    推荐理由:官方给出 1.3B 小模型处理高分辨率图像的编码成本与吞吐数据,可供端侧多模态选型参考。

  2. Hugging Face Blog63

    IBM 发布 Granite Embedding Multilingual R2 多语言嵌入模型,支持 32K 上下文

    IBM 发布两个 Apache 2.0 多语言嵌入模型 granite-embedding-97m-multilingual-r2 和 granite-embedding-311m-multilingual-r2,基于 ModernBERT,覆盖 200+ 语言、32K token 上下文,并支持 9 种编程语言的代码检索。

    推荐理由:两个模型以 Apache 2.0 覆盖 200+ 语言和 32K 上下文,读者可对照基准表判断小模型在多语言检索上的实际取舍。

5月9日周六
  1. Hugging Face Blog65

    AI2 发布 EMO:预训练混合专家实现涌现模块化

    AI2 发布 EMO,一个 1B 激活、14B 总参数(8 专家激活、128 专家总量)的混合专家模型,在 1 万亿 token 上端到端预训练,让模块化结构直接从数据中涌现。保留 25% 专家时各基准平均性能只下降约 1%,仅保留 12.5% 专家时平均下降约 3%,而同等架构的标准 MoE 在相同子集设置下明显退化。团队同时开源了 EMO 模型、同等数据的标准 MoE 基线和训练代码。

    推荐理由:EMO 让专家子集按语义领域自然分组,只保留 12.5% 专家仍接近全模型性能,为稀疏 MoE 的部署取舍提供参考。

5月8日周五
  1. @diegocabezas0166

    OpenAI 将新语音模型 gpt-realtime-2 上线 API,开发者 @diegocabezas01 发布了自建演示。被引用的 @sama 表示,人们正越来越多地用语音与 AI 交互,尤其是在需要输入大量上下文的场景,并称这是相当大的一步进展,同时提到团队正在改进聊天中的语音体验。

    引用Sam Altman (@sama)@sama

    people are really starting to use voice to interact with AI, especially when they have a lot of context to dump. GPT-Realtime-2 comes to the API today; it is a pretty big step forward. (we are working on improvements to voice in chat.)

    推荐理由:OpenAI 把 gpt-realtime-2 语音模型带进 API,并附开发者自建演示,可据此了解语音交互的使用场景。

5月5日周二
  1. OpenAI News65

    OpenAI 将 GPT-5.5 Instant 更新为 ChatGPT 默认模型

    OpenAI 将 GPT-5.5 Instant 更新为 ChatGPT 的默认模型,称其回答更智能、更准确,并减少了幻觉。该版本同时改进了个性化控制。

    推荐理由:ChatGPT 默认模型换为 GPT-5.5 Instant,读者可据此了解回答准确度、幻觉与个性化控制的调整方向。

  2. OpenAI News75

    OpenAI 发布 GPT-5.5 Instant 系统卡

    OpenAI 发布 GPT-5.5 Instant 系统卡,说明这款最新 Instant 模型沿用系列安全缓解方案。这是首款在网络安全与生物化学预备类别中被按高能力模型处理并实施相应保障的 Instant 模型,基线对比对象为 GPT-5.3 Instant。卡片还注明不存在 GPT-5.4 Instant,并把 GPT-5.5 称为 GPT-5.5 Thinking 以避免混淆。

    推荐理由:系统卡交代 GPT-5.5 Instant 沿用系列安全缓解方案,并首次把 Instant 模型在网络安全与生物化学预备类别按高能力处理。

5月1日周五
  1. @CalmPromptsHQ69

    SenseTime 将 SenseNova U1 Lite 系列开源,该系列基于 NEO-unify 架构,原生统一多模态理解与生成,提供 8B 和 A3B 两个紧凑版本。官方称其在开源模型中效率领先,并用单一模型单流生成交错的文本与图像,面向指南、知识插图、海报、PPT、漫画等信息密集版式。模型权重已在 Hugging Face 和 GitHub 发布,并可自托管。

    引用SenseTime (@SenseTime_AI)@SenseTime_AI

    𝗦𝗲𝗻𝘀𝗲𝗡𝗼𝘃𝗮 𝗨1 𝗟𝗶𝘁𝗲 𝗦𝗲𝗿𝗶𝗲𝘀 𝗶𝘀 𝗻𝗼𝘄 𝗼𝗽𝗲𝗻 𝘀𝗼𝘂𝗿𝗰𝗲! Built on the 𝗡𝗘𝗢-𝘂𝗻𝗶𝗳𝘆 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲, it natively unifies multimodal understanding and generation, delivering: •𝗦𝗢𝗧𝗔 𝗘𝗳𝗳𝗶𝗰𝗶𝗲𝗻𝗰𝘆 𝗔𝗺𝗼𝗻𝗴 𝗢𝗽𝗲𝗻-𝗦𝗼𝘂𝗿𝗰𝗲 𝗠𝗼𝗱𝗲𝗹𝘀: Compact models (8B & A3B) delivering commercial-grade performance and exceptional cost efficiency. Leading performance among open-source models across a wide range of understanding, reasoning, and generation benchmarks. •𝗡𝗮𝘁𝗶𝘃𝗲 𝗜𝗺𝗮𝗴𝗲–𝗧𝗲𝘅𝘁 𝗜𝗻𝘁𝗲𝗿𝗹𝗲𝗮𝘃𝗲𝗱 𝗚𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝗼𝗻: Generate coherent interleaved text and images in a single flow using one model; ideal for practical applications like guides, where visuals turn complex information into intuitive insights. •𝗛𝗶𝗴𝗵-𝗗𝗲𝗻𝘀𝗶𝘁𝘆 𝗜𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻 𝗥𝗲𝗻𝗱𝗲𝗿𝗶𝗻𝗴: Strong capabilities in dense visual communication, generating richly structured layouts for knowledge illustrations, posters, PPTs, comics and other information-rich formats. 𝗛𝘂𝗴𝗴𝗶𝗻𝗴 𝗙𝗮𝗰𝗲: huggingface.co/collections/s… 𝗚𝗶𝘁𝗛𝘂𝗯: github.com/OpenSenseNova/Sen… 𝗗𝗶𝘀𝗰𝗼𝗿𝗱: discord.gg/cxkwXWjp  @huggingface @github

    推荐理由:SenseNova U1 Lite 以 8B、A3B 紧凑规模开源权重,读者可据此衡量自托管图文版式生成的成本门槛。