跳到正文

推理能力

模型推理能力的进展:思维链、推理模型、数学与逻辑基准的突破与争议。

当前仅显示精选新闻

最新精选

第 161–180 条 · 共 211 条
6月18日周四
  1. IT Home78

    华为昇腾 0 Day 支持智谱 GLM-5.2 模型,提供全面推理优化

    昇腾 0 Day 支持智谱 GLM-5.2,为编程与长程任务提供全面推理优化,A3 系列产品已支持单双机及大 EP 推理部署。官方围绕 MoE 大融合算子、通信与计算融合、多 Token 预测、高并发调度、智能缓存以及 PD 分离与 Prefix Cache 等关键技术做推理优化。GLM-5.2 于 6 月 17 日上线并开源,已在 Day 0 完成与华为昇腾等国产算力平台的推理适配。

    推荐理由:原文列出昇腾针对 GLM-5.2 的多项推理优化技术,读者可了解国产算力平台适配长上下文模型的工程路径。

6月17日周三
  1. Google Blog: AI69

    Google 在 Nature 发表 AMIE 研究:医疗 AI 从诊断对话扩展到长期疾病管理

    Google 在 Nature 发表研究,展示医疗 AI 系统 AMIE 从一次性诊断对话扩展到使用药物目录和临床指南的长期疾病管理,该系统基于 Gemini 模型的长上下文能力,包含共情对话智能体和深度思考的管理推理智能体。

    推荐理由:盲测中 AMIE 的整体疾病管理推理与 21 名全科医生持平,计划精确度和指南一致性更高。

6月11日周四
  1. Google DeepMind72

    Google DeepMind 发布开源实验模型 DiffusionGemma,文本生成速度最高提升 4 倍

    Google DeepMind 发布 Apache 2.0 开源的实验性文本扩散模型 DiffusionGemma,为 26B MoE 架构、推理时激活 3.8B 参数,在专用 GPU 上文本生成最高提速 4 倍,单张 NVIDIA H100 超过 1000 tokens/秒,量化后可放入 18GB 显存。

    推荐理由:官方给出了具体吞吐数字、显存占用和适用场景边界,读者可以据此判断扩散文本生成适合哪些本地工作流。

6月9日周二
6月5日周五
  1. Hugging Face Blog61

    NVIDIA 发布 Nemotron 3.5 Content Safety 多模态安全模型

    NVIDIA 发布 Nemotron 3.5 Content Safety,基于 Google Gemma 3 4B IT 微调,把多模态输入、自定义企业策略与可审计推理链统一到单次推理调用中,并保持 12 种语言显式训练和约 140 种语言的零样本泛化。

    推荐理由:相比 Nemotron 3,3.5 版把多模态审核、自定义策略与可审计推理链合并到一次调用中。

6月4日周四
  1. @kimmonismus81

    NVIDIA 发布 Nemotron 3 Ultra,一款完全开源的 550B MoE 模型,激活参数 55B,权重、训练数据与完整配方全部公开。该模型采用混合 Mamba-Attention MoE 架构,NVIDIA 称其在长输出智能体任务上的吞吐量约为同类开源模型的 6 倍,同时保持相同准确率。

    引用NVIDIA AI (@NVIDIAAI)@NVIDIAAI

    Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. Video

    推荐理由:原文给出 550B 开源模型的权重、训练数据与完整配方,并说明其在长任务智能体上的吞吐表现,读者可据此判断开源前沿模型的可复现程度。

  2. @MiniMax_AI69

    MiniMax 官方宣布 M3 在 1M token 下解码速度提升 15.6 倍,并感谢 Fireworks AI 为其提供推理支持,用户可直接试用。其引用的 Fireworks AI 内容显示,M3 采用 MiniMax Sparse Attention(MSA),模型权重发布后也将在 Fireworks 社区提供。

    引用Fireworks AI (@FireworksAI_HQ)@FireworksAI_HQ

    MiniMax M3 arrives with MiniMax Sparse Attention (MSA), 15.6x faster decoding at 1M tokens. We're partnering with @MiniMax_AI to power the inference behind this week's launch. Head to minimax.io to take it for a spin. Once the model weights are released, M3 will be available to the Fireworks community.

    推荐理由:官方给出 M3 在 1M token 下解码提速 15.6 倍,读者可据此判断其长上下文推理的工程取向。

6月3日周三
  1. @berryxia71

    微软 AI 在 Build 上发布七个全新 MAI 模型,官方称并非简单迭代,而是从零开始、干净数据血统、零蒸馏训练的一整个家族,涵盖推理、编码、图像、转录与语音并各有 Flash 版本。

    引用Microsoft AI (@MicrosoftAI)@MicrosoftAI

    Seven new models launching at Build: let’s go! Reasoning. Code. Image. Transcribe. Voice. Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models Thread 🧵 #MSBuild

    推荐理由:文中梳理了七个 MAI 模型的任务分工与基准数字,可据此了解微软从零训练、任务专精的模型家族路线。

  2. @eliebakouch73

    微软 MAI 技术报告因透明度受到讨论,报告显示该模型未使用合成数据或来自此前模型的蒸馏,推理、智能体行为与工具调用均在 post-training 阶段完整习得。报告给出模型各迭代阶段的精确 MFU 及对应变化,并公开完整 scaling ladder 配方,作者称这是他在同规模技术报告中第一次见到如此详细的披露。

    引用Mustafa Suleyman (@mustafasuleyman)@mustafasuleyman

    Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities. - It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks. - And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end. Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing. Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI. - Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost. All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat. Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost. Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare. Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: microsoft.ai/news/building-a…

    推荐理由:作者逐点点评微软 MAI 技术报告,读者可了解其无合成数据与蒸馏的训练取舍及 scaling ladder 的公开细节。

6月2日周二
6月1日周一
  1. @MiniMax_AI76

    MiniMax 宣布 M3 在 OrcaRouter 首日上线,首周使用可享 50% 折扣。引用信息显示,M3 支持最高 1M token 上下文(最低保证 512K),具备新一代稀疏注意力、更快推理和更强的智能体工作流,并称其为最受期待的开源模型发布之一。

    引用OrcaRouter 🐳 (@OrcaRouter)@OrcaRouter

    🚀 @MiniMax_AI M3 is now available on OrcaRouter. One of the most anticipated open model releases, bringing next-gen sparse attention, ultra-long context handling (up to 1M tokens of context with a guaranteed minimum of 512K), faster inference, and stronger agentic workflows. Try it out here: orcarouter.ai/models/minimax…

    推荐理由:MiniMax M3 首日上线 OrcaRouter,读者可据此了解其长上下文与智能体工作流的能力定位。

  2. MiniMax Blog81

    MiniMax 发布 M3:1M 上下文、原生多模态与前沿编码能力

    MiniMax 发布 M3 模型,采用自研稀疏注意力架构 MSA,支持最高 1M token 上下文,SWE-Bench Pro 得分 59.0%、Terminal-Bench 2.1 得分 66.0%,并原生支持图像与视频输入及桌面操作。

    推荐理由:官方详解了 MSA 架构与 1M 上下文的具体实现,并给出多项基准成绩和真实任务案例,可据此评估其编码与长程智能体能力。

5月29日周五
5月28日周四
5月27日周三
  1. @kimmonismus76

    小米 MiMo-V2.5 系列 API 定价永久下调,最高较此前降低 99%,并统一所有上下文长度的定价。MiMo Token 套餐在同价下可用 token 增加 5–8 倍,现有用户的 Token Plan credits 将全部重置,MiMo-V2.5-TTS 限期免费,新价格于 5 月 26 日 6:00 PM PDT 生效。作者 @kimmonismus 指出 MiMo 2.5 Pro 现在与 DeepSeek V4 Pro 价格相同。

    引用Xiaomi MiMo (@XiaomiMiMo)@XiaomiMiMo

    🚀 Better inference efficiency, lower costs, broader access. MiMo-V2.5 Series API pricing is now permanently reduced — by up to 99% compared to previous pricing. ✨ Unified pricing across all context lengths. MiMo Token Plans have also been upgraded: • 5–8× more usable tokens at the same price • Simpler and more transparent billing rules 🎁 As a thank-you to current users, all current Token Plan credits will be fully reset. 🎧 MiMo-V2.5-TTS remains free for a limited time. ⏰ Effective May 26 at 6:00 PM PDT. These improvements are powered by continued inference optimization and serving efficiency upgrades across the MiMo stack. 🛠️ We’ll also publish a detailed technical blog on the inference optimizations later — stay tuned.

    推荐理由:原文列出 API 价格最高降 99% 与 Token 套餐升级,读者可借此观察推理成本下探对模型选型的影响。