跳到正文

多模态

文本之外的能力:视觉理解、图文混合、音视频输入输出的模型与产品进展。

122条精选相关主题图像生成AI 视频语音与音频

最新精选

第 81–100 条 · 共 122 条
5月20日周三
  1. @berryxia72

    Gemini 3.5 Flash 已在 ZenMux 上线并提供免费试用,作者实测用它从提示词生成完整 HTML 递归树生长动画,全程耗时 77.56 秒。该模型在 MCP Atlas、Toolathlon、Finance Agent 等榜单拿下第一,MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,并兼容主流 API 格式。

    推荐理由:作者用递归树动画实测 Gemini 3.5 Flash 的生成速度,并列出其在 Agent 榜单与多模态基准上的成绩。

  2. @berryxia79

    Google DeepMind 发布 Gemini 3.5 Flash,Artificial Analysis 预发布测试显示其 Intelligence Index 得 55 分,比 Gemini 3 Flash 高 9 分。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in its Flash family, which has traditionally has offered faster, lower-cost alternatives to Gemini Pro models. Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash, driven primarily by agentic performance gains and hallucination reduction. It achieves speeds of over 280 output tokens/s, but higher token usage and token pricing make it over 5x more costly to run the Intelligence Index than Gemini 3 Flash, and 75% more costly than Gemini 3.1 Pro. Gemini 3.5 Flash is $1.50/1M input and $9/1M output tokens, Gemini 3 Flash was $0.5/$3 per 1M input/output tokens, a 3x increase. The rest of the increase was driven by higher token usage when running our benchmarks Key results for Gemini 3.5 Flash with ‘high’ thinking level: ➤ 9 point Intelligence Index improvement: Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash. This places it ahead of Grok 4.3 (high, 53) and Claude Sonnet 4.6 (max, 52). The model improves across nearly all evaluations, with the largest gains coming from agentic evaluations and AA-Omniscience (knowledge and hallucination). On AA-Omniscience, Gemini 3.5 Flash improves by 11 points, driven primarily by reduced hallucinations, with its hallucination rate falling to 61%, a 31 point decrease compared to Gemini 3 Flash ➤ Agentic capability improvements: Gemini 3.5 Flash improves substantially over Gemini 3 Flash across our agentic evaluations, in both GDPval-AA (real-world agentic tasks) and Tau2-Bench Telecom (agentic tool use). Its GDPval-AA result is especially notable, achieving an Elo of 1656, well ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), and just behind GPT-5.4 (xhigh, 1674). This represents a meaningful step forward for Google in agentic performance, which has historically been a relative weakness for Gemini models ➤ Speed-intelligence frontier: Gemini 3.5 Flash achieves speeds of over 280 output tokens per second, ~70% faster than Gemini 3 Flash and models such as gpt-oss-120b and GPT-5.4 mini (xhigh). With its 55 Intelligence Index score, this places Gemini 3.5 Flash on the speed-intelligence Pareto frontier alongside Gemini 3.1 Pro and Gemini 3.1 Flash-Lite, reinforcing Google’s strength in models balancing speed and intelligence ➤ 5.5x increase in cost to run: Gemini 3.5 Flash costs $1,552 to run the Artificial Analysis Intelligence Index, 5.5x more than Gemini 3 Flash and 75% more than Gemini 3.1 Pro. This is driven by increases in both token usage and token prices. Output token usage is broadly unchanged from Gemini 3 Flash (73M vs. 72M), but input token usage increases significantly, driven primarily by an increase in the number of turns in agentic evaluations. Gemini 3.5 Flash is priced 3x higher than Gemini 3 Flash at $1.50/$9.00 per 1M input/output tokens, with a 90% discount for cached input tokens ➤ Google continues to lead multimodal performance: Gemini 3.5 Flash is multimodal, supporting image, video, and speech input alongside text. This differs from many proprietary models, including Claude Opus 4.7, Grok 4.3, and GPT-5.5, which support image input only. In our multimodal evaluation, MMMU-Pro, Gemini 3.5 Flash scores 84% - the highest score recorded. This puts models from Google in the top two spots, with Gemini 3.1 Pro scoring 82% Key model details: ➤ Context window: Retains the same 1M context window as Gemini 3 Flash ➤ Multimodality: Text, image, video and speech input with text output only ➤ Pricing: $1.50/$9.00 per million input/output tokens, with a 90% discount for cached input tokens Congratulations @GoogleDeepMind , @sundarpichai and @demishassabis on the great release!

    推荐理由:借 Artificial Analysis 的预发布基准,可以看到 Gemini 3.5 Flash 在智能与速度上的提升及其成本代价。

  3. @berryxia75

    Google DeepMind 发布 Gemini Omni,将 Gemini 的智能与生成媒体系统融合,可先定义角色再放入任意场景并保持外貌、动作和光影一致,也支持用自然语言改风格、加效果或重拍已有视频。

    引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMind

    We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 Video

    推荐理由:Gemini Omni 把生成视频做成可对话编辑的对象,并同步在 Gemini App 等入口上线,读者可据此观察视频生成向可编辑素材演进。

  4. @berryxia73

    Gemini Omni 开始向全球 Google AI Plus、Pro 和 Ultra 订阅用户推出,首先支持视频输出。它不只构建看起来真实的场景,还能推理接下来应该发生什么,将对物理学的直观理解与 Gemini 对历史、科学和文化背景的知识结合起来。

    推荐理由:材料交代了 Gemini Omni 面向订阅层的开放节奏与视频优先的输出形态,读者可据此判断上手门槛。

  5. @minchoi81

    Google 发布新模型 Gemini Omni,可从任意输入创建内容,首发支持视频,被形容为视频版 Nano Banana。该模型已在 Gemini App、Flow 和 YouTube 上线,API 支持即将推出。

    引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganK

    Introducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video

    推荐理由:原文给出 Gemini Omni 从任意输入生成视频的能力,以及 Gemini App、Flow 和 YouTube 的上线入口。

  6. @GeminiApp75

    Gemini Omni 今天登陆 Gemini 应用,面向付费订阅用户开放。该功能支持文本、图像和视频的任意组合输入,用户可在 Gemini 中附加相册里的视频并对其进行修改。

    引用Google Gemini (@GeminiApp)@GeminiApp

    Gemini Omni is coming to the Gemini app for paid subscribers today. It lets you bring your ideas to life using any combination of text, images, and video inputs. Just open up Gemini, attach a video from your camera roll, and change it around. It’s that simple. #GoogleIO

    推荐理由:Gemini Omni 面向付费订阅用户开放,给出文本、图像与视频混合输入的用法,读者可据此判断可用范围。

  7. @kimmonismus82

    Gemini Omni 发布,官方介绍这是一款可从任意输入生成内容的新模型,首批聚焦视频,已在 Gemini App、Flow 和 YouTube 上线,API 支持即将推出。转发的作者 @kimmonismus 称这是真正的惊艳时刻,认为它是迈向 AGI 的世界模型,能从任何输入生成任何内容。

    引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganK

    Introducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video

    推荐理由:官方说明了 Gemini Omni 从任意输入生成内容的能力和首批视频场景,读者可据此了解当前可用渠道与范围。

  8. @GeminiApp66

    Gemini app 新增 SynthID 验证,覆盖图像、视频和音频,用户可直接提问内容是否由 AI 生成。该应用同时加入 C2PA Content Credentials 验证,用于判断内容是否为相机拍摄的未修改原图,或已被哪些工具修改。两项功能今日起在 Gemini app 逐步推送。

    推荐理由:Gemini app 把 AI 生成内容核验做成了对话式查询,读者可了解多模态内容溯源在应用内的落地方式。

  9. @Google78

    Google 在 GeminiApp 上线 Gemini Omni,支持任意组合文本与视觉输入生成电影感视频,即日起面向全球 Google AI Plus、Pro 和 Ultra 订阅用户开放。它提供对话式的视频创建与编辑方式,用户用简单提示词即可实现电影级变焦或更换背景,无需昂贵设备或技术术语。

    推荐理由:原文给出 Gemini Omni 的输入方式与开放订阅档位,读者可据此判断视频创作门槛的变化。

5月19日周二
  1. @berryxia66

    Odyssey 发布 Agora-1,一个多智能体世界模型,人类与 AI 可同时进入同一模拟世界并实时互动、互相影响。官方推出可游玩的研究预览,用 Agora-1 模拟多人 GoldenEye 死亡竞赛,模型实时生成画面和声音,整个世界持续更新。

    引用Odyssey (@odysseyml)@odysseyml

    Introducing Agora-1, a multi-agent world model. Multiple participants—human or AI—can now interact inside the same world simulation, all in real-time. Try our playable research preview today, with Agora-1 simulating a multiplayer GoldenEye deathmatch! Video

    推荐理由:世界模型从单人视频生成扩展到多人实时共享模拟,读者可据此了解人机共处同一模拟世界的当前形态。