推荐理由:作者用递归树动画实测 Gemini 3.5 Flash 的生成速度,并列出其在 Agent 榜单与多模态基准上的成绩。
多模态
全部主题文本之外的能力:视觉理解、图文混合、音视频输入输出的模型与产品进展。
最新精选
第 81–100 条 · 共 122 条@berryxia@berryxia精选AI 评分7272 
@berryxia@berryxia精选AI 评分6565 
推荐理由:作者用同一递归树 Prompt 记录生成耗时与效果,并列出 Agent 与多模态榜单成绩,可快速判断该模型的能力定位。
@berryxia@berryxia精选AI 评分6666 Gemini 3.5 Flash 已在 ZenMux 上线并提供免费额度体验,也可通过 API 调用。作者用它跑递归二叉树生长测试,从输入提示词到生成完整 HTML 动画网页耗时 77.56 秒。

推荐理由:材料给出同一提示词下的生成耗时与多项榜单成绩,可作为了解 Gemini 3.5 Flash 速度与 Agent 能力的参考。
@berryxia@berryxia精选AI 评分7979
引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlysGoogle’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in its Flash family, which has traditionally has offered faster, lower-cost alternatives to Gemini Pro models. Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash, driven primarily by agentic performance gains and hallucination reduction. It achieves speeds of over 280 output tokens/s, but higher token usage and token pricing make it over 5x more costly to run the Intelligence Index than Gemini 3 Flash, and 75% more costly than Gemini 3.1 Pro. Gemini 3.5 Flash is $1.50/1M input and $9/1M output tokens, Gemini 3 Flash was $0.5/$3 per 1M input/output tokens, a 3x increase. The rest of the increase was driven by higher token usage when running our benchmarks Key results for Gemini 3.5 Flash with ‘high’ thinking level: ➤ 9 point Intelligence Index improvement: Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash. This places it ahead of Grok 4.3 (high, 53) and Claude Sonnet 4.6 (max, 52). The model improves across nearly all evaluations, with the largest gains coming from agentic evaluations and AA-Omniscience (knowledge and hallucination). On AA-Omniscience, Gemini 3.5 Flash improves by 11 points, driven primarily by reduced hallucinations, with its hallucination rate falling to 61%, a 31 point decrease compared to Gemini 3 Flash ➤ Agentic capability improvements: Gemini 3.5 Flash improves substantially over Gemini 3 Flash across our agentic evaluations, in both GDPval-AA (real-world agentic tasks) and Tau2-Bench Telecom (agentic tool use). Its GDPval-AA result is especially notable, achieving an Elo of 1656, well ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), and just behind GPT-5.4 (xhigh, 1674). This represents a meaningful step forward for Google in agentic performance, which has historically been a relative weakness for Gemini models ➤ Speed-intelligence frontier: Gemini 3.5 Flash achieves speeds of over 280 output tokens per second, ~70% faster than Gemini 3 Flash and models such as gpt-oss-120b and GPT-5.4 mini (xhigh). With its 55 Intelligence Index score, this places Gemini 3.5 Flash on the speed-intelligence Pareto frontier alongside Gemini 3.1 Pro and Gemini 3.1 Flash-Lite, reinforcing Google’s strength in models balancing speed and intelligence ➤ 5.5x increase in cost to run: Gemini 3.5 Flash costs $1,552 to run the Artificial Analysis Intelligence Index, 5.5x more than Gemini 3 Flash and 75% more than Gemini 3.1 Pro. This is driven by increases in both token usage and token prices. Output token usage is broadly unchanged from Gemini 3 Flash (73M vs. 72M), but input token usage increases significantly, driven primarily by an increase in the number of turns in agentic evaluations. Gemini 3.5 Flash is priced 3x higher than Gemini 3 Flash at $1.50/$9.00 per 1M input/output tokens, with a 90% discount for cached input tokens ➤ Google continues to lead multimodal performance: Gemini 3.5 Flash is multimodal, supporting image, video, and speech input alongside text. This differs from many proprietary models, including Claude Opus 4.7, Grok 4.3, and GPT-5.5, which support image input only. In our multimodal evaluation, MMMU-Pro, Gemini 3.5 Flash scores 84% - the highest score recorded. This puts models from Google in the top two spots, with Gemini 3.1 Pro scoring 82% Key model details: ➤ Context window: Retains the same 1M context window as Gemini 3 Flash ➤ Multimodality: Text, image, video and speech input with text output only ➤ Pricing: $1.50/$9.00 per million input/output tokens, with a 90% discount for cached input tokens Congratulations @GoogleDeepMind , @sundarpichai and @demishassabis on the great release!
推荐理由:借 Artificial Analysis 的预发布基准,可以看到 Gemini 3.5 Flash 在智能与速度上的提升及其成本代价。
@berryxia@berryxia精选AI 评分7676 
推荐理由:归纳了 Gemini 3.5 系列、Omni 世界模型与硬件落地的要点,可据此了解 Google 在智能体方向的推进节奏。
@berryxia@berryxia精选AI 评分7575
引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMindWe’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 Video
推荐理由:Gemini Omni 把生成视频做成可对话编辑的对象,并同步在 Gemini App 等入口上线,读者可据此观察视频生成向可编辑素材演进。
@berryxia@berryxia精选AI 评分7373 
推荐理由:材料交代了 Gemini Omni 面向订阅层的开放节奏与视频优先的输出形态,读者可据此判断上手门槛。
@minchoi@minchoi精选AI 评分8181
引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganKIntroducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video
推荐理由:原文给出 Gemini Omni 从任意输入生成视频的能力,以及 Gemini App、Flow 和 YouTube 的上线入口。
微信公众号(Mp2RSS 合集)精选AI 评分8989 Google I/O 2026 开发者大会回顾,Gemini 3.5 Flash 与 Antigravity 2.0 等集中发布
Google 在 I/O 2026 开发者大会上集中发布了 Gemini 3.5 Flash、Gemini Omni Flash、Antigravity 2.0、Gemini Spark 等模型与产品更新。
推荐理由:逐项梳理 Google I/O 2026 的模型与 Agent 更新,读者可据此了解 Google 这半年的产品布局。
@GeminiApp@geminiapp精选AI 评分7070 推荐理由:官方公布了 Gemini Omni 面向全球订阅用户的开放范围,以及后续将支持的图像、音频输出格式。
@GeminiApp@geminiapp精选AI 评分6666 推荐理由:官方给出 Gemini Omni 的多模态输入与视频生成能力,读者可据此判断视频生成工作流的入口变化。
@OpenRouter@openrouter精选AI 评分8080 
推荐理由:原文给出 Gemini 3.5 Flash 的定价与上下文规格,并对比其在编码和工具调用上优于 Gemini 3.1 Pro。
@GeminiApp@geminiapp精选AI 评分7575 Gemini Omni 今天登陆 Gemini 应用,面向付费订阅用户开放。该功能支持文本、图像和视频的任意组合输入,用户可在 Gemini 中附加相册里的视频并对其进行修改。
引用Google Gemini (@GeminiApp)@GeminiAppGemini Omni is coming to the Gemini app for paid subscribers today. It lets you bring your ideas to life using any combination of text, images, and video inputs. Just open up Gemini, attach a video from your camera roll, and change it around. It’s that simple. #GoogleIO
推荐理由:Gemini Omni 面向付费订阅用户开放,给出文本、图像与视频混合输入的用法,读者可据此判断可用范围。
@kimmonismus@kimmonismus精选AI 评分8282 引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganKIntroducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video
推荐理由:官方说明了 Gemini Omni 从任意输入生成内容的能力和首批视频场景,读者可据此了解当前可用渠道与范围。
@GeminiApp@geminiapp精选AI 评分6666 
推荐理由:Gemini app 把 AI 生成内容核验做成了对话式查询,读者可了解多模态内容溯源在应用内的落地方式。
@GeminiApp@geminiapp精选AI 评分6969 
推荐理由:Gemini 应用开放 Gemini Omni 视频创作,读者可从覆盖的订阅档位和入口判断自身是否可用。
@Google@google精选AI 评分7878 
推荐理由:原文给出 Gemini Omni 的输入方式与开放订阅档位,读者可据此判断视频创作门槛的变化。
IT Home精选AI 评分7676 谷歌搜索迎 25 年来最大改版,AI 重塑搜索体验与交互方式
谷歌在 2026 谷歌 I/O 开发者大会上宣布谷歌搜索迎来 25 年来最大改版,用 AI 全面重塑搜索入口与交互方式,由 Gemini 3.5 Flash 模型提供快速响应。
推荐理由:改版把 AI Mode、多模态输入与后台智能体一并纳入搜索入口,呈现搜索交互从关键词到完整需求的转向。
@Google@google精选AI 评分6868 Google 的 Gemini Omni 支持用对话编辑视频,用户可在多轮提示中重新构想动作、改变视角或调整光线。每次指令都会基于上一条,让角色保持一致、物理表现成立,并让场景记住先前内容。

推荐理由:原文给出多轮对话编辑视频的交互方式,读者可以了解角色一致性和物理表现在连续指令下的处理。
@berryxia@berryxia精选AI 评分6666
引用Odyssey (@odysseyml)@odysseymlIntroducing Agora-1, a multi-agent world model. Multiple participants—human or AI—can now interact inside the same world simulation, all in real-time. Try our playable research preview today, with Agora-1 simulating a multiplayer GoldenEye deathmatch! Video
推荐理由:世界模型从单人视频生成扩展到多人实时共享模拟,读者可据此了解人机共处同一模拟世界的当前形态。