华为昇腾 0day 适配 Kimi K3,支持训练与推理部署
月之暗面于 2026 年 7 月 27 日正式开源 Kimi K3,总参数量达 2.8 万亿。华为宣布昇腾实现 0day 极速适配,训练侧基于 MindSpeed MM 在 Atlas 800 A3、Atlas 900 A3 SuperPoD 上完成减层训练基础功能适配。
推荐理由:原文给出昇腾对 Kimi K3 的训练与推理适配细节,读者可了解国产算力与开源大模型的协同路径。
月之暗面 Kimi 系列模型与产品动态:K 系列开源模型、长上下文技术与产品演进的追踪。
当前仅显示精选新闻月之暗面于 2026 年 7 月 27 日正式开源 Kimi K3,总参数量达 2.8 万亿。华为宣布昇腾实现 0day 极速适配,训练侧基于 MindSpeed MM 在 Atlas 800 A3、Atlas 900 A3 SuperPoD 上完成减层训练基础功能适配。
推荐理由:原文给出昇腾对 Kimi K3 的训练与推理适配细节,读者可了解国产算力与开源大模型的协同路径。
Kimi K3 发布开放权重并公开 47 页技术报告,包含 2.8T 总参数、104B 激活参数和 100 万 token 上下文。报告重点在 KDA、AttnRes 与 Stable LatentMoE 三项架构改动,分别处理模型变长、变深、变宽后的信息流问题。
推荐理由:技术报告披露的 KDA、AttnRes 与 LatentMoE 三项架构改动,可帮读者理解长上下文和长程 Agent 训练的工程取舍。
月之暗面公开 Kimi K3 的完整权重与 47 页技术报告,报告称相较 K2 整体 scaling 效率提升约 2.5 倍。技术报告显示,K3 采用 3:1 的 KDA 线性注意力与 MLA 混合架构、AttnRes 层间检索,以及 896 个专家、每个 token 激活 16 个的 MoE 配置。
推荐理由:结合 K3 技术报告的关键架构与训练改动,对比 DeepSeek 的路线差异,呈现开源模型在预训练一侧的进展。
月之暗面于 7 月 27 日晚开放 Kimi K3 的模型权重、技术报告和关键 Infra 技术,Anthropic 联合创始人兼 CEO Dario Amodei 随后发文回应称,Anthropic 从未主张禁止开放权重模型。
推荐理由:材料把 K3 开源所需的千万级硬件门槛与开放权重政策争论并置,读者可据此判断开源模型的实际使用边界。
月之暗面正式开源 Kimi K3,参数量 2.8 万亿、激活参数 1040 亿,支持 100 万 token 上下文并具备原生视觉理解,同步发布技术报告。
推荐理由:技术报告公开了注意力、跨层残差与 MoE 负载均衡的具体改动,可借此看清单个开源模型在万亿参数规模上的工程取舍。
月之暗面开源 Kimi K3 的模型权重和技术报告,2.8 万亿参数 MoE,具备原生视觉理解能力,支持 100 万 token 上下文窗口。
推荐理由:权重、技术报告与训练 Infra 一并公开,呈现 2.8 万亿参数 MoE 从架构到线上服务的工程链路。
月之暗面开源 Kimi K3 模型,参数规模 2.8 万亿,同步放出模型权重、技术报告以及 MoonEP、FlashKDA、AgentEnv 三项支撑训练的 Infra 技术。
推荐理由:2.8 万亿参数开源模型把权重、技术报告与训练 Infra 一并放出,可对照开源模型的规模上限与工程路径。
SGLang 和 Miles 发布对 Kimi K3 的 Day-0 支持,分别覆盖推理与 RL 训练。K3 为 2.8T 参数的 3 万亿参数级首个开源模型,采用 69 层 KDA 线性注意力与 24 层 MLA 的混合架构和 1M token 上下文窗口。
推荐理由:原文由 SGLang 团队自述 Day-0 支持的实现细节,给出内存管理、投机解码与并行策略的具体方案和实测数字,可迁移到异构注意力架构的部署。
Moonshot AI 于7月16日发布旗舰模型 Kimi K3,为 2.8T 参数 MoE 模型,权重定于7月27日开放,是 DeepSeek R1 之后离前沿最近的开源模型。作者分析其多项榜单表现、中国实验室的资本效率优势、习近平在 WAIC 表态支持开源,并讨论开源模型对封闭实验室的经济冲击及政策风险。
推荐理由:作者走访过月之暗面团队,结合榜单和政策动向分析开源前沿模型的格局变化,给出资本效率与开放封闭平衡的独特视角。
月之暗面新一代旗舰模型Kimi K3上线48小时后,请求量逼近现有集群承载极限,官方即日起暂停C端新用户订阅,已订阅用户可继续使用。Kimi K3为全球首个开源的3万亿参数级模型,在前端代码竞技场以1679分排名第一,SWE Marathon拿到42.0分。据彭博社,K3发布以来月之暗面日销售额至少增长6倍,6月年度经常性收入已达3亿美元,公司正推进赴港上市,最快可能6个月内完成IPO。
推荐理由:K3上线48小时即因请求量触顶而暂停C端新订阅,可见爆款模型对上算力与会员体系的即时压力。
Kimi K3 发布,成为全球第一个开源的3T级别大模型,在 Frontend Code Arena 排名第一并大幅领先 Fable 5。其编程能力在 DeepSWE 拿到 67.5 分,SWE Marathon 42.0 分为所有模型最高,48 小时自主完成一颗 45nm 芯片的设计优化与验证,并从零构建了名为 MiniTriton 的 GPU 编译器。
推荐理由:原文给出前端、编程与芯片设计等多维实测对比,读者可据此判断这款开源大模型的性价比与能力边界。
月之暗面发布旗舰模型 Kimi K3,在 Arena 前端代码/Web Dev 榜单以 1679 分登顶,领先 Fable 5 48 分、高于 GPT-5.6 Sol 61 分。
推荐理由:Kimi K3 登顶 Arena 前端榜单并提价近 4 倍,读者可对照其稀疏 MoE 架构与真实任务成本的差异。
月之暗面发布 Kimi K3,这是一个总参数 2.8T 的 MoE 开源模型,官方称其整体仍落后于 Claude Fable 5 和 GPT-5.6 Sol,但稳定超过其他所有模型。
推荐理由:官方给出了与 Fable 5、GPT-5.6 Sol 的评测对比和定价,读者可据此判断开源模型当前的位置与使用成本。
月之暗面于 7 月 16 日晚上线 Kimi K3,为 2.8 万亿参数 MoE 模型,原生支持视觉理解和 100 万 token 上下文,官方称整体扩展效率较 K2 提升约 2.5 倍。
推荐理由:作者以四项自测呈现 K3 的能力与短板,并对照定价与上下文分层的变化说明开源旗舰模型的使用门槛。
Kimi K3 正式发布,上下文窗口 1M、总参数 2.8T,官方在基准测试中于 14 项里的 11 项击败 GPT-5.6 Sol,全部 14 项击败 Opus 4.8,并在 Frontend Code Arena 以 1679 分位列第一,超越 Claude Fable 5。
推荐理由:官方跑分与第三方评测同时给出对比数据,读者可据此判断 K3 在前端能力和成本上相对同级模型的位置。
Kimi K3 发布并已在网页端上线,参数量达2.8万亿,是目前公开的开源模型中规模最大的一个。它基于Kimi Delta Attention(KDA)混合线性注意力机制,并引入Attention Residuals结构,原生支持视觉理解,上下文窗口达100万token。
推荐理由:原文列出Kimi K3的参数规模、KDA与AttnRes架构改动及多项榜单跑分,可据此判断当前开源模型的能力位置。
Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-recording-to-code, and more 🔹Plugins for stocks, financial reports, academic papers, with more coming 🔹Supports the ACP protocol, and works with JetBrains, Zed, and more 🔹Hooks for custom tools and workflows Try it with Kimi K2.6 👉 kimi.com/code Issues, plugin ideas, and PRs welcome! Community feedback helps shape what ships next.🚀
推荐理由:原文列出升级后的零配置安装、视频上下文与插件能力,读者可据此判断编码智能体的门槛变化。
Meet Kimi Work - a local AI agent on your desktop that does the work for you. 🔹Native agent swarm: Up to 300 AI agents running in parallel on your local machine. 🔹Browser use: Paired with WebBridge extension, your agent will navigate websites in your browser: search, scroll, click, type and complete tasks. 🔹Built for Finance: Native global market data tool call from Yahoo Finance and World Bank - no complex API setup required. 🔹Memory system: Kimi Desktop keeps a running diary of your preferences, past decisions, and context to know you better. Available for macOS (Apple Silicon) and Windows. 🔗Try it now: kimi.com/products/kimi-work Video
推荐理由:官方披露了本地代理规模、浏览器操作与财经数据原生调用等能力,读者可据此判断桌面端智能体的落地形态。
Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past releases @cursor_ai has released Composer 2.5, the latest model in its Composer line. Composer 2.5 scored 62 on our Coding Agent Index, a 14 point gain over Composer 2 (48). This puts it in third place of our tested agents, behind only Claude Opus 4.7 (max) in Claude Code (66) and GPT-5.5 (xhigh reasoning) in Codex (65). These cost $4.10 and $4.82 per task respectively, ~10x the cost of Composer 2.5 Fast ($0.44) and ~60x the cost of Composer 2.5 standard ($0.07). Key results for Composer 2.5 in Cursor CLI: ➤ Cost-quality Pareto frontier: At $0.07 (standard) and $0.44 (Fast) per task, Composer 2.5 is cheaper than every other agent scoring above 60 on the Index. Medium-effort peers cost $1.24–$2.21 per task; higher-effort variants land 3-4 points above at $4.10–$4.82 ➤ Per-benchmark gains vs Composer 2: +35 points on SWE-Bench-Pro-Hard-AA (12% → 47%), +2 points on Terminal-Bench v2 (64% → 66%), and +3 points on SWE-Atlas-QnA (69% → 72%). At 47%, Composer 2.5's score on SWE-Bench-Pro-Hard-AA is comparable to Claude Opus 4.7 (max) in Claude Code ➤ Among the fastest coding agents: Composer 2.5 Fast runs at an average wall time of 6.7 minutes per task, the third-fastest agent on the Artificial Analysis Coding Agent Index, behind only Claude Opus 4.7 (medium) in Claude Code (5.8m) and GPT-5.5 (medium) in Cursor CLI (6.2m) ➤ Fast mode enables better responsiveness at 6x pricing: Fast runs 30% faster than standard Composer 2.5, but is ~6x the cost per task ($0.44 vs $0.07). Token pricing is 6x higher for Fast: $3.00/$15.00 vs $0.50/$2.50 per million input/output tokens Model details: ➤ Base model: Continued training on @Kimi_Moonshot's open weights Kimi K2.5 as with Composer 2, with Cursor reporting ~85% of total compute from its own additional training and reinforcement learning ➤ Pricing: $0.50/$2.50 per million input/output tokens for the standard variant; $3.00/$15.00 for the Fast variant (the default in Cursor) ➤ Available exclusively in Cursor: both Cursor IDE and Cursor CLI, an externally accessible API is not available Congratulations @cursor_ai and @mntruell on the impressive release!
推荐理由:推文用每任务成本对比 Composer 2.5 与两个更高分编码智能体,读者可据此权衡编码任务上的性能与花费。