跳到正文

模型发布

新模型的发布、开源与迭代:大模型厂商的旗舰更新、开源权重放出、性能与价格变化的第一时间记录。

当前仅显示精选新闻
575条精选相关主题产品更新论文研究开源生态

最新精选

第 221–240 条 · 共 575 条
8月24日周一
  1. @alibaba_cloud74

    Kimi-K3 已在阿里云 Model Studio 和 Qwen Cloud 上线,拥有 2.8T 参数和原生 1M token 上下文,面向仓库级软件工程、视觉创作以及能自主规划下一步的智能体。阿里云同步开放 1M token 免费试用,提供 1M-token free trial 与上线即有的 autoTPM 弹性容量。

    推荐理由:阿里云给出 Kimi-K3 的参数量、上下文规格与两个平台的试用入口,可据此判断其可用渠道与定位。

8月22日周六
  1. @AYi_AInotes68

    Google 发布 Gemini 3.7 Flash,官方将其定位为迄今最智能的工作马模型,Google 高管发推称这是其史上增长最快的一次模型发布。

    引用@OfficialLoganK@OfficialLoganK

    Gemini 3.7 is our fastest growing model launch to date, amazing to see the reception!!!

    推荐理由:文章给出 Gemini 3.7 Flash 的代码评测与价格变化,并讨论低价模型对 Agent 算力成本结构的影响。

8月21日周五
  1. @omarsar067

    DeepSeek-V4-Flash-Vision-Exp 已在 DeepSeek API 平台上线,该实验性多模态模型在文本能力上与 DeepSeek-V4-Flash 持平,在多模态智能体基准上大幅超越 V4-Flash 并接近 Opus-4.8。DeepSeek Harness 0.1.1 同日发布,内置对新模型的直接支持,调用模型名为 deepseek-v4-flash-vision-exp。作者另提示关注支持文本、图像和视频输入的 Ox Alpha 1M token 上下文,称其在编码和智能体任务上表现突出。

    引用@deepseek_ai@deepseek_ai

    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

    推荐理由:官方基准对比显示该实验模型在多模态智能体任务上接近 Opus-4.8,读者可据此了解多模态智能体的当前水平。

  2. @AYi_AInotes72

    DeepSeek 在 API 平台上线实验性多模态模型 deepseek-v4-flash-vision-exp,官方称其文本能力与 DeepSeek-V4-Flash 对齐,多模态 Agent 性能接近 Opus-4.8,在 Agents Last Exam 和 ZeroBench 上局部超过。

    原始视频预览图;未保存可播放视频URL
    引用@deepseek_ai@deepseek_ai

    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

    推荐理由:逐条列出这个实验性多模态模型的单图 token 上限、分级定价与免费 Files API 额度,可供评估视觉 Agent 的成本与接入方式。

  3. @kimmonismus67

    DeepSeek 在 API 平台上线实验性多模态模型 DeepSeek-V4-Flash-Vision-Exp,文本能力与 DeepSeek-V4-Flash 持平,多模态智能体基准表现接近 Opus-4.8。官方称该模型在多模态智能体基准上较 V4-Flash 大幅提升,可通过 model=deepseek-v4-flash-vision-exp 调用,DeepSeek Harness 0.1.1 同日发布并原生支持该模型。转发该消息的作者提到,取得这一成绩的是 Flash 系列中的小模型。

    引用@deepseek_ai@deepseek_ai

    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

    推荐理由:基准对比显示 Flash 小模型在多模态智能体评测上已接近 Opus-4.8,可供判断轻量模型的能力边界。

  4. @testingcatalog71

    DeepSeek-V4-Flash-Vision-Exp 多模态模型已在 DeepSeek API 平台上线,官方称其性能接近 Opus 4.8,文本能力与 DeepSeek-V4-Flash 持平。该模型支持 Chat Completions、Messages 和 Responses,可混合输入文本与图像,图像可通过 base64、外部 URL 或 Files API 提供,按 V4-Flash 价格计费,每张图片最多 384 tokens。随附基准表显示其在 Terminal Bench 2.1 得分 83.9、ApexBench 得分 36.5。

    引用@deepseek_ai@deepseek_ai

    Multimodal API support 🔌 🔹 Set model='deepseek-v4-flash-vision-exp' 🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing 🔹 Supports Chat Completions, Messages & Responses 🔹 Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API. Docs: https://t.co/USZ5gZ3wWB 3/n

    推荐理由:材料给出新模型与 Opus 4.8 及自家文本版本的基准对比和计费细节,便于判断实际可用性。

8月19日周三
  1. @thexpin68

    智谱 AI 发布新基础模型 GLM-5.3 的 API,聚焦复杂编码、网络安全防御和长程任务,在 Artificial Analysis Intelligence Index 得分 60,进入全球前沿模型梯队。该模型已接入 ZCode 和 GLM Coding Plan,与 Kimi K3 并列为领先开源模型,权重定于下周开源。

    推荐理由:原文交代了 GLM-5.3 的能力侧重、基准分数和开放节奏,读者可以据此判断其在开源与闭源模型中的位置。

8月18日周二
8月17日周一
  1. xAI News72

    xAI 发布 Grok 4.5,主打编码与智能体任务

    xAI 发布 Grok 4.5,称其为自家最强模型,面向编码、智能体任务和知识工作,并与 Cursor 一同训练。官方给出的成绩包括 SWE Marathon 通过率 29.0%、DeepSWE 1.0 得分 62.0%,并称其平均每个 SWE Bench Pro 任务输出 15,954 个 token,约为 Opus 4.8 (max) 的 4.2 分之一。

    推荐理由:原文给出多项基准对比、定价与 token 效率数据,读者可据此比较它与同代编码模型的位置。

  2. xAI News77

    xAI 发布 Grok 4 Fast,统一推理与非推理模式并支持 2M token 上下文

    xAI 发布 Grok 4 Fast,称其用平均少 40% 的思考 token 在基准上达到与 Grok 4 相当的表现,同一前沿基准性能的价格降低 98%。该模型采用统一架构,由系统提示词切换长链推理与快速响应,具备 2M token 上下文窗口和端到端训练的工具调用与网络、X 搜索能力,在 LMArena Search Arena 以 1163 Elo 排名第一。

    推荐理由:原文给出与 Grok 4 的基准对比、token 消耗和 API 定价,读者可据此判断推理模型的成本效率走向。

  3. xAI News66

    xAI 通过 API 开放编码模型 grok-build-0.1 公测

    xAI 通过 API 以公测形式开放编码模型 grok-build-0.1,该模型专门针对包括网页开发、调试和 MCP 支持在内的智能体编码任务训练,也是驱动 Grok Build 的同一模型。

    推荐理由:模型给出 100+ tokens/秒的服务速度与每百万 token 1 美元输入、2 美元输出的定价,便于评估智能体编码场景的成本与速度取舍。