跳到正文

多模态

文本之外的能力:视觉理解、图文混合、音视频输入输出的模型与产品进展。

当前仅显示精选新闻
238条精选相关主题图像生成AI 视频语音与音频

最新精选

第 181–200 条 · 共 238 条
6月1日周一
  1. @MiniMax_AI75

    MiniMax-M3 已在 Venice(@AskVenice)上线,作者称其开放权重,具备前沿编码与智能体能力,支持 1M 上下文和原生多模态,首日即可用并支持匿名访问。引用文案称其为首个兼具前沿编码与智能体性能并原生支持多模态的开放权重模型,面向长周期复杂工作流。

    引用Venice (@AskVenice)@AskVenice

    MiniMax-M3 by @MiniMax_AI is now live on Venice. The first open-weight model to deliver frontier coding and agentic performance with native multimodality. Built for long-horizon, complex workflows. Available anonymously.

    推荐理由:MiniMax-M3 以开放权重同时具备前沿编码与智能体能力,支持 1M 上下文和原生多模态,面向长周期复杂工作流。

  2. @MiniMax_AI71

    MiniMax 发布 M3 模型,并在发布当天同步上线 OpenRouter。该模型具备 1M token 上下文窗口、前沿编码与智能体能力,以及原生多模态,首周提供 50% 折扣。OpenRouter 称其为前沿级开源权重模型,原生多模态覆盖图像与视频。

    引用OpenRouter (@OpenRouter)@OpenRouter

    MiniMax-M3 is live on OpenRouter! A frontier-class open-weight model that combines a 1M-token context window, frontier coding and agentic performance, and native multimodality (image & video) in one model.

    推荐理由:MiniMax-M3 上线 OpenRouter 并给出首周折扣,可了解其 1M 上下文与多模态能力的组合方式。

  3. @MiniMax_AI66

    MiniMax 发布 MiniMax M3,称其为首个同时结合编码与 Agent、长上下文、原生多模态三项前沿能力的开源权重模型。编码与 Agentic 基准成绩为 SWE-Bench Pro 59.0%、Terminal Bench 2.1 66.0%、SWE-fficiency 34.8%、KernelBench Hard 28.8%、MCP Atlas 74.2%;MiniMax Sparse Attention 将上下文扩展至 1M,模型从 Step Zero 起原生多模态。API 已在 platform.minimax.io 上线,同时推出 MiniMax Code,权重与技术报告约 10 天后发布。

    推荐理由:MiniMax M3 放出编码、1M 上下文与多模态的基准成绩,可供开源权重模型选型时对照参考。

  4. MiniMax Blog81

    MiniMax 发布 M3:1M 上下文、原生多模态与前沿编码能力

    MiniMax 发布 M3 模型,采用自研稀疏注意力架构 MSA,支持最高 1M token 上下文,SWE-Bench Pro 得分 59.0%、Terminal-Bench 2.1 得分 66.0%,并原生支持图像与视频输入及桌面操作。

    推荐理由:官方详解了 MSA 架构与 1M 上下文的具体实现,并给出多项基准成绩和真实任务案例,可据此评估其编码与长程智能体能力。

5月29日周五
  1. @OpenRouter69

    阶跃星辰发布 Step 3.7 Flash,主打 agent 效率,采用 198B 稀疏 MoE、约 11B 激活、256K 上下文和 3 级推理,以 Apache 2.0 开放权重。该模型在 ClawEval-1.1 得 67.1、SimpleVQA Search 得 79.2 均列第一,SWE-PRO 得 56.3,并支持 Claude Code、MCP 等工具调用,可在 Mac Studio M4 Max、DGX Spark 等设备本地运行。

    引用StepFun (@StepFun_ai)@StepFun_ai

    ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SWE-PRO (56.3), 95.3 on V* Python. Open weights under Apache 2.0. Built for agentic, coding, search, and multimodal workflows — balancing speed, cost, and reliable execution. - 400 TPS. 198B sparse MoE, ~11B active. 256K context, 3 reasoning levels. - Understands UIs, charts, docs, images — then writes code or calls tools to act on what it sees. - Web + visual search reaches further: more sources, deeper follow-up. - Reliable tool use — less drift, fewer broken toolcalls. 98%+ on τ²-bench across all difficulty levels. - Works with Claude Code, KiloCode, Hermes Agent, OpenClaw, and protocols like MCP. - Runs locally on Mac Studio M4 Max, DGX Spark, AMD AI Max+ 395. GitHub: github.com/stepfun-ai/Step-3… HuggingFace: huggingface.co/stepfun-ai/St… GGUF: huggingface.co/stepfun-ai/St… ModelScope: modelscope.cn/models/stepfun… API: platform.stepfun.ai Blog: static.stepfun.com/blog/step…

    推荐理由:从开放权重和速度成本平衡切入 agent 场景,读者可对照其编程与搜索评测及本地部署支持。

5月28日周四
5月27日周三
  1. inclusionAI Hugging Face models68

    inclusionAI 发布 LLaDA2.0-Uni,统一多模态理解与生成的扩散 MoE 大模型

    inclusionAI 发布基于 MoE 的扩散大语言模型 LLaDA2.0-Uni,在单一模型中统一多模态理解与生成,支持文生图、图像理解、图像编辑以及交错生成与推理。

    推荐理由:模型把多模态理解与生成统一进扩散 LLM 框架,并公开 8 步解码与加速方案,读者可据此判断部署门槛。

5月22日周五
  1. Mistral AI69

    Mistral 发布 Mistral Medium 3.5 并在 Vibe 和 Le Chat 推出云端远程智能体

    Mistral 发布 Mistral Medium 3.5,一个 128B 稠密开源权重模型(modified MIT 许可),256k 上下文窗口,SWE-Bench Verified 得分 77.6%,τ³-Telecom 得分 91.4,自托管最少只需四块 GPU,API 定价每百万输入 token $1.5、输出 token $7.5。

    推荐理由:官方同步发布模型与云端异步智能体,给出基准分数、定价和开源权重,可对照评估其编码与长程任务能力。

5月20日周三
  1. @berryxia72

    Gemini 3.5 Flash 已在 ZenMux 上线并提供免费试用,作者实测用它从提示词生成完整 HTML 递归树生长动画,全程耗时 77.56 秒。该模型在 MCP Atlas、Toolathlon、Finance Agent 等榜单拿下第一,MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,并兼容主流 API 格式。

    推荐理由:作者用递归树动画实测 Gemini 3.5 Flash 的生成速度,并列出其在 Agent 榜单与多模态基准上的成绩。

  2. @berryxia79

    Google DeepMind 发布 Gemini 3.5 Flash,Artificial Analysis 预发布测试显示其 Intelligence Index 得 55 分,比 Gemini 3 Flash 高 9 分。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in its Flash family, which has traditionally has offered faster, lower-cost alternatives to Gemini Pro models. Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash, driven primarily by agentic performance gains and hallucination reduction. It achieves speeds of over 280 output tokens/s, but higher token usage and token pricing make it over 5x more costly to run the Intelligence Index than Gemini 3 Flash, and 75% more costly than Gemini 3.1 Pro. Gemini 3.5 Flash is $1.50/1M input and $9/1M output tokens, Gemini 3 Flash was $0.5/$3 per 1M input/output tokens, a 3x increase. The rest of the increase was driven by higher token usage when running our benchmarks Key results for Gemini 3.5 Flash with ‘high’ thinking level: ➤ 9 point Intelligence Index improvement: Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash. This places it ahead of Grok 4.3 (high, 53) and Claude Sonnet 4.6 (max, 52). The model improves across nearly all evaluations, with the largest gains coming from agentic evaluations and AA-Omniscience (knowledge and hallucination). On AA-Omniscience, Gemini 3.5 Flash improves by 11 points, driven primarily by reduced hallucinations, with its hallucination rate falling to 61%, a 31 point decrease compared to Gemini 3 Flash ➤ Agentic capability improvements: Gemini 3.5 Flash improves substantially over Gemini 3 Flash across our agentic evaluations, in both GDPval-AA (real-world agentic tasks) and Tau2-Bench Telecom (agentic tool use). Its GDPval-AA result is especially notable, achieving an Elo of 1656, well ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), and just behind GPT-5.4 (xhigh, 1674). This represents a meaningful step forward for Google in agentic performance, which has historically been a relative weakness for Gemini models ➤ Speed-intelligence frontier: Gemini 3.5 Flash achieves speeds of over 280 output tokens per second, ~70% faster than Gemini 3 Flash and models such as gpt-oss-120b and GPT-5.4 mini (xhigh). With its 55 Intelligence Index score, this places Gemini 3.5 Flash on the speed-intelligence Pareto frontier alongside Gemini 3.1 Pro and Gemini 3.1 Flash-Lite, reinforcing Google’s strength in models balancing speed and intelligence ➤ 5.5x increase in cost to run: Gemini 3.5 Flash costs $1,552 to run the Artificial Analysis Intelligence Index, 5.5x more than Gemini 3 Flash and 75% more than Gemini 3.1 Pro. This is driven by increases in both token usage and token prices. Output token usage is broadly unchanged from Gemini 3 Flash (73M vs. 72M), but input token usage increases significantly, driven primarily by an increase in the number of turns in agentic evaluations. Gemini 3.5 Flash is priced 3x higher than Gemini 3 Flash at $1.50/$9.00 per 1M input/output tokens, with a 90% discount for cached input tokens ➤ Google continues to lead multimodal performance: Gemini 3.5 Flash is multimodal, supporting image, video, and speech input alongside text. This differs from many proprietary models, including Claude Opus 4.7, Grok 4.3, and GPT-5.5, which support image input only. In our multimodal evaluation, MMMU-Pro, Gemini 3.5 Flash scores 84% - the highest score recorded. This puts models from Google in the top two spots, with Gemini 3.1 Pro scoring 82% Key model details: ➤ Context window: Retains the same 1M context window as Gemini 3 Flash ➤ Multimodality: Text, image, video and speech input with text output only ➤ Pricing: $1.50/$9.00 per million input/output tokens, with a 90% discount for cached input tokens Congratulations @GoogleDeepMind , @sundarpichai and @demishassabis on the great release!

    推荐理由:借 Artificial Analysis 的预发布基准,可以看到 Gemini 3.5 Flash 在智能与速度上的提升及其成本代价。

  3. @berryxia75

    Google DeepMind 发布 Gemini Omni,将 Gemini 的智能与生成媒体系统融合,可先定义角色再放入任意场景并保持外貌、动作和光影一致,也支持用自然语言改风格、加效果或重拍已有视频。

    引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMind

    We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 Video

    推荐理由:Gemini Omni 把生成视频做成可对话编辑的对象,并同步在 Gemini App 等入口上线,读者可据此观察视频生成向可编辑素材演进。

  4. @berryxia73

    Gemini Omni 开始向全球 Google AI Plus、Pro 和 Ultra 订阅用户推出,首先支持视频输出。它不只构建看起来真实的场景,还能推理接下来应该发生什么,将对物理学的直观理解与 Gemini 对历史、科学和文化背景的知识结合起来。

    推荐理由:材料交代了 Gemini Omni 面向订阅层的开放节奏与视频优先的输出形态,读者可据此判断上手门槛。

  5. @minchoi81

    Google 发布新模型 Gemini Omni,可从任意输入创建内容,首发支持视频,被形容为视频版 Nano Banana。该模型已在 Gemini App、Flow 和 YouTube 上线,API 支持即将推出。

    引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganK

    Introducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video

    推荐理由:原文给出 Gemini Omni 从任意输入生成视频的能力,以及 Gemini App、Flow 和 YouTube 的上线入口。