跳到正文

Google / Gemini

Google 与 DeepMind 的 AI 动态:Gemini 系列、Veo 视频模型、研究成果与产品生态的持续追踪。

当前仅显示精选新闻

最新精选

第 41–60 条 · 共 247 条
9月4日周五
  1. @sundarpichai87

    谷歌 CEO Sundar Pichai 公开祝贺 NVIDIA 收购 Hugging Face,并表示谷歌曾是 Hugging Face 的投资方,今后仍将继续保持合作,认为这会强化开源模型生态。他引用的 Jensen Huang 帖文称,NVIDIA 将成为 Hugging Face、其社区以及开放模型未来的好归宿,并提到开放模型有助于增强安全与网络安全、加速创新扩散并支持主权能力。

    引用@JensenHuang@JensenHuang

    Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 https://t.co/q8Om2Xc5ye

    推荐理由:Pichai 的祝贺透露谷歌曾是 Hugging Face 投资方并将继续合作,为这笔交易补上生态与股东视角。

  2. Google Research64

    Google Research 联合 HHMI Janelia 发布完整雄性果蝇大脑连接组

    Google Research 与 HHMI Janelia、MRC 分子生物学实验室及剑桥大学等合作,在 Cell 发表雄性果蝇大脑和中央神经系统的完整连接组,包含超 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。

    推荐理由:由参与方介绍完整雄性果蝇连接组,读者可了解其规模、AI 重建方法及对后续脑映射研究的用途。

9月3日周四
  1. Google DeepMind72

    Google DeepMind 发布 WeatherNext 3 全球 AI 天气模型

    Google DeepMind 与 Google Research 发布 WeatherNext 3,称其为最先进准确的全球 AI 天气模型。模型直接学习实时卫星数据和气象站观测,逐小时产出最高 5 公里分辨率预报,精度约为 WeatherNext 2 的五倍;中期降水预报 CRPS 相对 IMERG 提升最高 60%。

    推荐理由:官方介绍了新模型的分辨率、逐小时更新和降水精度提升,并说明其在 Google 各产品中的落地方式,读者可评估对天气相关业务的应用价值。

  2. @AYi_AInotes65

    Google 发布 Gemini 3.8 Flash,官方称其在智能体和编码能力上继续提升,是 6 周内第三个更新的 Flash 模型。作者对比称其价格约为 Claude Opus 5 的 1.5 折,在法律 10.0% 对 6.7%、长视频 87.8% 对 75.4% 上反超;但 OSWorld 电脑操作 59.0% 对 75.4%、通用 Agent 规划 19.1% 对 51.8% 仍落后。

    引用@OfficialLoganK@OfficialLoganK

    Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think! https://t.co/Cj07lCBtp8

    推荐理由:文中列出 Gemini 3.8 Flash 与 Opus 5 的多项基准和价格对比,读者可据此看到廉价模型与旗舰模型当下的分工边界。

  3. @kimmonismus67

    前沿模型竞争格局在很短时间内从 OpenAI 与 Anthropic 双强之争,扩展为 OpenAI、Anthropic、xAI、Meta 多方并跑、Google 重新加入的局面,中国开源权重模型也紧随其后。

    引用@ArtificialAnlys@ArtificialAnlys

    Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code

    推荐理由:作者把前沿模型竞争格局的变化讲清楚,并指出中国开源权重模型已贴近第一梯队,可与2025年的撞墙争论对照。

  4. @fofrAI74

    Google DeepMind 发布两款 Gemini 新模型:3.8 Flash 与 3.8 Flash Cyber。3.8 Flash 在软件工程、智能体任务和多步推理上较 3.7 Flash 有明显提升,3.8 Flash Cyber 则主打前沿水平的漏洞检测与自动修补。

    引用@GoogleDeepMind@GoogleDeepMind

    Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level vulnerability detection and automated patching.

    推荐理由:Google 一次发布两款 Gemini 3.8 Flash 模型,分别面向智能体任务与代码安全,可对照 3.7 Flash 看提升方向。

9月2日周三
  1. @omarsar067

    Google DeepMind 发布 Gemini 3.8 Flash 和 Gemini 3.8 Flash Cyber 两款模型。前者在软件工程、智能体任务和多步推理上较 3.7 Flash 有明显提升,后者面向网络安全,具备前沿级漏洞检测与自动修补能力。Elvis Saravia 称 Gemini 3.8 Flash 的定价有吸引力,并认为 3.8 Flash Cyber 在修补能力上处于 Pareto 前沿。

    引用@GoogleDeepMind@GoogleDeepMind

    Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level vulnerability detection and automated patching.

    推荐理由:随推附上的对比表把 Gemini 3.8 Flash 的定价与多项基准成绩与 Claude、GPT-5.6 并列,便于横向比较。

  2. @OfficialLoganK65

    Gemini 3.8 Flash 发布,主打智能体与编码能力提升,是 6 周内第三个更新的 Flash 模型。推文附带的对比表显示,其输入价格 $0.75/1M tokens、输出价格 $3.75/1M tokens,Terminal-bench 2.1 得 89.4%,LBVBench 长视频理解得 87.8%(agentic)。优惠价有效期至 2026 年 12 月 31 日,2027 年 1 月 1 日起将调整为输入 $1.50/1M tokens、输出 $7.50/1M tokens。

    推荐理由:原文列出 Gemini 3.8 Flash 的价格与多项基准对比,读者可据此判断其智能体和编码能力相较前代及竞品的位置。

  3. @GoogleAI74

    Google 将 3.8 Flash 开放给 Google AI Pro 与 Ultra 订阅者,覆盖 Gemini App、Google 搜索 AI Mode 和 Google Sheets。开发者可通过 Google AI Studio 与 Android Studio 的 Gemini API 使用,在 antigravity 中探索 agent-first 工作流,并在 stitch 中用 3.8 Flash 生成 UI。Gemini Enterprise 用户可在下拉模型菜单或 Gemini Enterprise Agent Platform 中选择该模型。https://t.co/OzPm3Et10i

    推荐理由:原文列出 3.8 Flash 在订阅端、开发者工具和企业平台的具体开放入口,读者可据此判断自己能否用上。

  4. @GoogleDeepMind67

    Google DeepMind 发布 Gemini 3.8 Flash 和 3.8 Flash Cyber 两款新模型,前者面向智能体与软件工程,后者面向代码安全。官方称 3.8 Flash 是迄今最智能的模型,在软件工程、智能体任务和多步推理上相较 3.7 Flash 有明显提升。3.8 Flash Cyber 则具备前沿水平的漏洞检测与自动化补丁修复能力。

    原始视频预览图;未保存可播放视频URL

    推荐理由:官方列出两款 3.8 Flash 的定位差异与相对 3.7 Flash 的提升方向,便于判断编码与安全场景的选型。

  5. @GoogleDeepMind65

    Google DeepMind 宣布推出 Fairwind 计划,为政府和可信合作伙伴提供 3.8 Flash Cyber 的访问权限,用于保护关键基础设施和国家安全。Gemini 3.8 Flash 正在 Antigravity 中推送,并通过 Google AI Studio 和 Android Studio 的 API 提供。Google AI Pro 和 Ultra 订阅者可在 Gemini App 及 Google Search 的 AI Mode 中使用 3.8 Flash。

    原始视频预览图;未保存可播放视频URL

    推荐理由:官方公布了 Fairwind 计划与 Gemini 3.8 Flash 的推送渠道,读者可据此了解面向政府和关键基础设施的接入方式。

  6. Google Developers Blog62

    Google 复盘 AI Agents Challenge 最强提交背后的 4 种工程模式

    Google for Startups AI Agents Challenge 评选结束后,官方从高分提交中总结出四种工程模式:双向 MCP、事件驱动并发、同标准降级(Gemini 3.1 Pro 503 时回退 Gemini 3.6 Flash 并共用同一校验函数)、以及模型调用前的分层路由。

    推荐理由:文章从真实参赛代码中提炼四个可复用的工程模式,覆盖工具暴露、并发、降级校验和分层路由,可直接迁移到自己的智能体项目。

  7. @GoogleDeepMind67

    Gemini 的智能体视频理解开始在 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 上通过 Google AI Studio 的 API 推送,并即将登陆 Gemini App。该能力不再扫描整个文件,而是跨视频字幕、音频和画面做推理,动态调整帧率以定位所需片段,长视频内容(从 10 分钟指南到数小时录像)的效率提升最明显。官方图表显示,在 1H-VideoQA 与 LVBench 等长视频基准上,单次查询 token 从约 30 万至 40 万降至 5 万以内,准确率同步小幅上升。

    推荐理由:官方图表给出长视频场景 token 消耗大幅下降而准确率上升的对比,读者可据此判断长视频处理的成本空间。

  8. Google DeepMind68

    Google DeepMind 为 Gemini 推出 agentic video understanding 视频分析功能

    Google DeepMind 在 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 上推出 agentic video understanding,动态检索视频片段以替代固定帧率处理,token 消耗最多降 88%,成本最多降 66%,准确率最多提升 7%。

    推荐理由:官方给出 token、成本与准确率的具体降幅和开启方式,长视频处理场景下可评估是否切换处理模式。

8月28日周五