跳到正文

模型发布

新模型的发布、开源与迭代:大模型厂商的旗舰更新、开源权重放出、性能与价格变化的第一时间记录。

当前仅显示精选新闻
530条精选相关主题产品更新论文研究开源生态

最新精选

第 441–460 条 · 共 530 条
5月22日周五
  1. @berryxia68

    Stable Audio 3 发布,可在 Mac 本地运行音乐生成,在 M5 Pro 上达到 59x realtime。其 Sm 模式更快、Medium 模式质量更高,LoRA 微调不到 1 小时完成。官方提供 MLX 优化版,可用一行命令安装。

    引用dadabots (@dadabots)@dadabots

    🥳 Announcing Stable Audio 3 🍕 🏆 fastest music models ever 💻 runs on MacBookPro M-series 🧪 break it plz 🧠 LoRA finetune in < 1h 📷 Sm = faster, Medium = qualityer ⚡ 59x realtime on M5 Pro One-liner fast install: curl -LsSf dadabots.com/_/sa3-mac | bash Video

    推荐理由:Stable Audio 3 给出在 Mac 本地运行音乐生成的安装方式与性能数据,可据此判断本地音乐生成工作流的可行性。

  2. @alibaba_cloud73

    阿里云 Qwen3.7-Max 已在 OpenRouter 上线。OpenRouter 称其为 Qwen3.7 系列旗舰,面向以智能体为核心的工作,涵盖编码、办公与生产力任务以及长程自主执行,在编码和智能体基准上较 Qwen3.6 有明显提升,并支持显式 prompt caching。

    引用OpenRouter (@OpenRouter)@OpenRouter

    The new Qwen3.7-Max from @Alibaba_Qwen is live on OpenRouter. The flagship of the Qwen3.7 series, built for agent-centric work: coding, office and productivity tasks, and long-horizon autonomous execution. Big jumps in coding and agent benchmarks over Qwen3.6, with explicit prompt caching for repeated context.

    推荐理由:Qwen3.7-Max 已在 OpenRouter 上线,面向编码与办公的智能体场景,读者可了解这一旗舰版本的定位。

  3. @OpenRouter69

    阿里发布最新旗舰模型 Qwen3.7-Max,定位为面向 Agent 场景的基础模型。官方称其可完成前端原型、多文件重构与调试等端到端编码任务,通过 MCP 集成和多智能体编排充当办公助手,并曾在内核优化任务上连续运行 35 小时、完成 1000+ 次工具调用。模型支持 Claude Code、OpenClaw、Qwen Code 等不同脚手架,API 已在 Alibaba Model Studio 上线,也可在 Qwen Studio 试用。

    引用Qwen (@Alibaba_Qwen)@Alibaba_Qwen

    📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get things done: 🧑‍💻 Coding agent, end to end. Frontend prototypes, multi-file refactors, real debugging — nails it. 🗂️ A reliable office and productivity assistant. Get your work done through MCP integrations and multi-agent orchestration. ⏱️ Long-horizon autonomy. 35 hours straight on a kernel optimization task — 1,000+ tool calls, zero hand-holding. 🔌 Scaffold-agnostic. Claude Code, OpenClaw, Qwen Code, or your own stack. Consistent reliability everywhere. API's up on Alibaba Model Studio. You can also take it for a spin on Qwen Studio. Go build something wild!🏃🏃‍♂️ 📖 Blog: qwen.ai/blog?id=qwen3.7 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.7… ⚡️ API:modelstudio.console.alibabac…

    推荐理由:原文列出端到端编码、MCP 集成与长时间自主运行等能力,可用以判断该旗舰模型在 Agent 场景中的定位。

5月21日周四
  1. @kimmonismus72

    阿里发布旗舰模型 Qwen3.7-Max,面向智能体场景,官方称其在一次内核优化任务中自主运行 35 小时、发起 1,158 次工具调用,并在单个注意力内核上取得 10 倍加速,模型已上线阿里云 Model Studio,也可在 Qwen Studio 试用。

    引用Qwen (@Alibaba_Qwen)@Alibaba_Qwen

    📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get things done: 🧑‍💻 Coding agent, end to end. Frontend prototypes, multi-file refactors, real debugging — nails it. 🗂️ A reliable office and productivity assistant. Get your work done through MCP integrations and multi-agent orchestration. ⏱️ Long-horizon autonomy. 35 hours straight on a kernel optimization task — 1,000+ tool calls, zero hand-holding. 🔌 Scaffold-agnostic. Claude Code, OpenClaw, Qwen Code, or your own stack. Consistent reliability everywhere. API's up on Alibaba Model Studio. You can also take it for a spin on Qwen Studio. Go build something wild!🏃🏃‍♂️ 📖 Blog: qwen.ai/blog?id=qwen3.7 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.7… ⚡️ API:modelstudio.console.alibabac…

    推荐理由:作者把 35 小时自主优化的传播印象与实际范围区分开,并单独讨论智能体能力泛化这一论断。

  2. @alibaba_cloud66

    阿里云披露 Qwen3.7 的自主进化实验:在约 35 小时连续自主执行中,模型完成 432 次 kernel 评估、跨 1158 次工具调用,独立编写、编译、剖析并迭代改进 Extend Attention Kernel。在多种工作负载下,该 kernel 相对 Triton 参考实现取得 10.0x 几何平均加速,更多细节见 qwen.ai/blog?id=qwen3.7。

    推荐理由:官方披露 Qwen3.7 连续自主执行约 35 小时优化注意力 kernel 的过程,可作为观察自主编码智能体能力的参照。

  3. @alibaba_cloud70

    阿里云发布旗舰模型 Qwen3.7-Max,定位为面向智能体时代的基础模型,API 已在 Model Studio 上线。官方称其可端到端完成编码任务,包括前端原型、多文件重构与调试,并可通过 MCP 集成和多智能体编排承担办公与生产力助手工作。

    推荐理由:官方列出编码、长任务与多种脚手架兼容能力,读者可据此判断这款旗舰模型在智能体工作流中的定位。

5月20日周三
  1. @berryxia72

    Gemini 3.5 Flash 已在 ZenMux 上线并提供免费试用,作者实测用它从提示词生成完整 HTML 递归树生长动画,全程耗时 77.56 秒。该模型在 MCP Atlas、Toolathlon、Finance Agent 等榜单拿下第一,MMMU-Pro 83.6%、CharXiv Reasoning 84.2%,并兼容主流 API 格式。

    推荐理由:作者用递归树动画实测 Gemini 3.5 Flash 的生成速度,并列出其在 Agent 榜单与多模态基准上的成绩。

  2. @berryxia79

    Google DeepMind 发布 Gemini 3.5 Flash,Artificial Analysis 预发布测试显示其 Intelligence Index 得 55 分,比 Gemini 3 Flash 高 9 分。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in its Flash family, which has traditionally has offered faster, lower-cost alternatives to Gemini Pro models. Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash, driven primarily by agentic performance gains and hallucination reduction. It achieves speeds of over 280 output tokens/s, but higher token usage and token pricing make it over 5x more costly to run the Intelligence Index than Gemini 3 Flash, and 75% more costly than Gemini 3.1 Pro. Gemini 3.5 Flash is $1.50/1M input and $9/1M output tokens, Gemini 3 Flash was $0.5/$3 per 1M input/output tokens, a 3x increase. The rest of the increase was driven by higher token usage when running our benchmarks Key results for Gemini 3.5 Flash with ‘high’ thinking level: ➤ 9 point Intelligence Index improvement: Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash. This places it ahead of Grok 4.3 (high, 53) and Claude Sonnet 4.6 (max, 52). The model improves across nearly all evaluations, with the largest gains coming from agentic evaluations and AA-Omniscience (knowledge and hallucination). On AA-Omniscience, Gemini 3.5 Flash improves by 11 points, driven primarily by reduced hallucinations, with its hallucination rate falling to 61%, a 31 point decrease compared to Gemini 3 Flash ➤ Agentic capability improvements: Gemini 3.5 Flash improves substantially over Gemini 3 Flash across our agentic evaluations, in both GDPval-AA (real-world agentic tasks) and Tau2-Bench Telecom (agentic tool use). Its GDPval-AA result is especially notable, achieving an Elo of 1656, well ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), and just behind GPT-5.4 (xhigh, 1674). This represents a meaningful step forward for Google in agentic performance, which has historically been a relative weakness for Gemini models ➤ Speed-intelligence frontier: Gemini 3.5 Flash achieves speeds of over 280 output tokens per second, ~70% faster than Gemini 3 Flash and models such as gpt-oss-120b and GPT-5.4 mini (xhigh). With its 55 Intelligence Index score, this places Gemini 3.5 Flash on the speed-intelligence Pareto frontier alongside Gemini 3.1 Pro and Gemini 3.1 Flash-Lite, reinforcing Google’s strength in models balancing speed and intelligence ➤ 5.5x increase in cost to run: Gemini 3.5 Flash costs $1,552 to run the Artificial Analysis Intelligence Index, 5.5x more than Gemini 3 Flash and 75% more than Gemini 3.1 Pro. This is driven by increases in both token usage and token prices. Output token usage is broadly unchanged from Gemini 3 Flash (73M vs. 72M), but input token usage increases significantly, driven primarily by an increase in the number of turns in agentic evaluations. Gemini 3.5 Flash is priced 3x higher than Gemini 3 Flash at $1.50/$9.00 per 1M input/output tokens, with a 90% discount for cached input tokens ➤ Google continues to lead multimodal performance: Gemini 3.5 Flash is multimodal, supporting image, video, and speech input alongside text. This differs from many proprietary models, including Claude Opus 4.7, Grok 4.3, and GPT-5.5, which support image input only. In our multimodal evaluation, MMMU-Pro, Gemini 3.5 Flash scores 84% - the highest score recorded. This puts models from Google in the top two spots, with Gemini 3.1 Pro scoring 82% Key model details: ➤ Context window: Retains the same 1M context window as Gemini 3 Flash ➤ Multimodality: Text, image, video and speech input with text output only ➤ Pricing: $1.50/$9.00 per million input/output tokens, with a 90% discount for cached input tokens Congratulations @GoogleDeepMind , @sundarpichai and @demishassabis on the great release!

    推荐理由:借 Artificial Analysis 的预发布基准,可以看到 Gemini 3.5 Flash 在智能与速度上的提升及其成本代价。

  3. @berryxia75

    Google DeepMind 发布 Gemini Omni,将 Gemini 的智能与生成媒体系统融合,可先定义角色再放入任意场景并保持外貌、动作和光影一致,也支持用自然语言改风格、加效果或重拍已有视频。

    引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMind

    We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 Video

    推荐理由:Gemini Omni 把生成视频做成可对话编辑的对象,并同步在 Gemini App 等入口上线,读者可据此观察视频生成向可编辑素材演进。

  4. @minchoi81

    Google 发布新模型 Gemini Omni,可从任意输入创建内容,首发支持视频,被形容为视频版 Nano Banana。该模型已在 Gemini App、Flow 和 YouTube 上线,API 支持即将推出。

    引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganK

    Introducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon! Video

    推荐理由:原文给出 Gemini Omni 从任意输入生成视频的能力,以及 Gemini App、Flow 和 YouTube 的上线入口。