跳到正文

数据与训练

训练侧的门道:数据集构建、合成数据、预训练与后训练方法、算力与训练成本。

当前仅显示精选新闻

最新精选

第 61–80 条 · 共 87 条
7月20日周一
  1. Nathan Lambert: Interconnects80

    Kimi K3 发布在即,Nathan Lambert 解析开源权重模型如何改写前沿格局

    Moonshot AI 于7月16日发布旗舰模型 Kimi K3,为 2.8T 参数 MoE 模型,权重定于7月27日开放,是 DeepSeek R1 之后离前沿最近的开源模型。作者分析其多项榜单表现、中国实验室的资本效率优势、习近平在 WAIC 表态支持开源,并讨论开源模型对封闭实验室的经济冲击及政策风险。

    推荐理由:作者走访过月之暗面团队,结合榜单和政策动向分析开源前沿模型的格局变化,给出资本效率与开放封闭平衡的独特视角。

7月17日周五
  1. Hugging Face Blog69

    NVIDIA NeMo Automodel 与 Diffusers 集成,可大规模微调视频和图像模型

    NVIDIA 和 Hugging Face 推出 NeMo Automodel 与 Diffusers 的集成,为 Hub 上任意 Diffusers 格式模型提供生产级分布式扩散训练,无需转换 checkpoint 或重写模型。

    推荐理由:NVIDIA 与 Hugging Face 把分布式扩散训练接入 Diffusers,无需转换 checkpoint 或重写模型,可直接微调现有开源扩散模型。

  2. 量子位 · 微信公众号78

    Kimi K3 发布:前端竞技场排名第一,48小时自主设计芯片并写出GPU编译器

    Kimi K3 发布,成为全球第一个开源的3T级别大模型,在 Frontend Code Arena 排名第一并大幅领先 Fable 5。其编程能力在 DeepSWE 拿到 67.5 分,SWE Marathon 42.0 分为所有模型最高,48 小时自主完成一颗 45nm 芯片的设计优化与验证,并从零构建了名为 MiniTriton 的 GPU 编译器。

    推荐理由:原文给出前端、编程与芯片设计等多维实测对比,读者可据此判断这款开源大模型的性价比与能力边界。

7月16日周四
7月9日周四
  1. Google Research61

    Google Research 发布可穿戴健康基础模型 SensorFM,基于超一万亿分钟传感器数据预训练

    Google Research 推出 SensorFM,从500万同意参与者采集的超一万亿分钟多模态可穿戴传感器数据中自监督预训练,覆盖 PPG、加速度、EDA、皮肤温度和测高五类模态。

    推荐理由:原文给出SensorFM的训练规模、跨35项健康任务的泛化结果和Agent集成实验,读者可据此评估可穿戴健康基础模型的实际能力。

6月26日周五
6月25日周四
  1. Hugging Face Blog65

    NVIDIA NeMo AutoModel 开源库发布,MoE 微调吞吐提升 3.4 至 3.7 倍

    NVIDIA 发布 NeMo AutoModel 开源库,基于 Transformers v5,在 Qwen3-30B-A3B 和 Nemotron 3 Nano 30B A3B 微调中实现 3.4 至 3.7 倍训练吞吐提升与 29% 至 32% 显存下降,只需改动一行 import。

    推荐理由:只需改动一行 import,MoE 微调吞吐即可提升 3.4 至 3.7 倍、显存降低 29% 至 32%,读者可据此判断迁移成本。

6月10日周三
  1. 量子位 · 微信公众号78

    Anthropic 新模型 Fable 5 被指护栏误触频繁,防蒸馏机制会静默降低回答质量

    Anthropic 今天凌晨发布 Fable 5 与 Mythos 5 后,多位用户实测发现 Fable 5 的安全护栏触发频率远高于官方宣称的不到 5%,普通编码任务或日常打招呼都可能被自动切回 Opus 4.8。

    推荐理由:文章梳理了 Fable 5 安全护栏与防蒸馏机制的设计细节,可帮助读者理解用户实测中被切回 Opus 4.8 的体感落差。

  2. @kimmonismus68

    Anthropic 的 Fable 5 保障机制在前沿 LLM 开发场景下不会直接拒绝或提醒用户,而是通过 prompt modification、steering vectors 和 PEFT 等方式悄悄降低模型自身的有效性。

    引用NomoreID (@Hangsiin)@Hangsiin

    When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model’s capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic.

    推荐理由:材料呈现模型在敏感领域被悄然降能的机制,可供观察厂商在能力与安全之间的取舍方式。

6月6日周六
  1. IT Home77

    谷歌每月向 SpaceX 支付 9.2 亿美元租用 AI 算力,含约 11 万张英伟达 GPU

    谷歌与 SpaceX 达成云计算合作,计划自 2026 年 10 月起至 2029 年 6 月每月支付 9.2 亿美元租用数据中心算力。租赁内容涵盖至少 11 万张英伟达 GPU、CPU 等芯片对应的计算能力,主要面向训练和推理等 AI 高密度场景。华尔街日报认为,这既能缓解谷歌的算力供应紧张与扩容周期压力,也为 SpaceX 的 AI 业务新增一条收入来源,为其 IPO 提供叙事筹码。

    推荐理由:协议披露了按月支付的算力租赁规模与起止时间,读者可据此观察谷歌的算力缺口与 SpaceX 的 AI 收入布局。

6月3日周三
  1. @eliebakouch73

    微软 MAI 技术报告因透明度受到讨论,报告显示该模型未使用合成数据或来自此前模型的蒸馏,推理、智能体行为与工具调用均在 post-training 阶段完整习得。报告给出模型各迭代阶段的精确 MFU 及对应变化,并公开完整 scaling ladder 配方,作者称这是他在同规模技术报告中第一次见到如此详细的披露。

    引用Mustafa Suleyman (@mustafasuleyman)@mustafasuleyman

    Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities. - It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks. - And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end. Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing. Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI. - Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost. All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat. Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost. Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare. Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: microsoft.ai/news/building-a…

    推荐理由:作者逐点点评微软 MAI 技术报告,读者可了解其无合成数据与蒸馏的训练取舍及 scaling ladder 的公开细节。

6月2日周二
6月1日周一
  1. Hugging Face Blog80

    NVIDIA 发布 Cosmos 3:首个面向物理 AI 推理与动作的开源全模态模型

    NVIDIA 在 Hugging Face 上发布 Cosmos 3,这是一个面向物理 AI 的开源全模态模型。它基于 Mixture-of-Transformers 架构,将世界生成、物理推理与动作生成整合到单一模型中。

    推荐理由:NVIDIA 将世界生成、物理推理与动作生成整合到一个模型中,并给出开源模型与 Diffusers 接入方式。

5月29日周五
  1. Hugging Face Blog62

    PyTorch 性能分析(一):torch.profiler 入门指南

    Hugging Face 博客发布 PyTorch profiling 系列第一篇,以矩阵乘法加加法为例讲解 torch.profiler 的用法。文章说明 profiler 会导出统计表格和 Perfetto 时序 trace 两类产物,并演示如何用 warmup 消除冷启动开销、区分 CPU 与 GPU lane 以及读懂嵌套的调度链,最后附上速查表。

    推荐理由:面向初学者的 torch.profiler 阅读指南,用两个矩阵算子串起从 CPU 调度到 CUDA 内核的完整链路,并附上速查表。

5月27日周三
  1. Hugging Face Blog66

    TRL 推出 Delta Weight Sync,异步 RL 每步权重传输从 1.2GB 降至 20-35MB

    Hugging Face 在 TRL 中实现 Delta Weight Sync,只把 RL 相邻优化步骤间真正变化的 bf16 权重编码为稀疏 safetensors 文件,上传到 Hugging Face Bucket 后由 vLLM 拉取。

    推荐理由:TRL 通过稀疏增量同步把异步 RL 每步权重传输量压缩约两个数量级,并允许 trainer 与推理集群分离部署。

5月21日周四
  1. Tomer Tunguz82

    SpaceX 递交 S-1,披露 Starlink、发射与 AI 三大业务数据

    SpaceX 递交 S-1,披露 2025 年 187 亿美元合并营收与 66 亿美元调整后 EBITDA。文件把公司分为 Space、Starlink 和 AI 三个分部,Starlink 贡献 61% 营收、2025 年运营利润 44 亿美元,AI 分部当年投入 64 亿美元建设 COLOSSUS 数据中心并训练 Grok。

    推荐理由:S-1 数据把 SpaceX 拆成卫星、发射与 AI 三块业务,读者可借此比较 AI 算力投入与收入回报的差距。

  2. Simon Willison89

    SpaceX S-1 披露 Anthropic 每月支付 12.5 亿美元租用算力

    SpaceX 提交的 S-1 文件披露,Anthropic 与其签订云服务协议,获得 COLOSSUS 和 COLOSSUS II 的算力访问权限,每月支付 12.5 亿美元至 2029 年 5 月,2026 年 5 月和 6 月按较低费用爬坡。协议任一方可提前 90 天通知终止。文件还提到 Grok 5 目前正在 COLOSSUS II 训练。

    推荐理由:S-1 文件披露出 Anthropic 与 SpaceX 算力协议的具体金额和期限,读者可借这份一手披露了解大模型训练算力的商业条款。

  3. @AYi_AInotes68

    阿易 AI Notes 引用泄露音频称,扎克伯格在 4 月 30 日全员会上表示 Meta 正用员工的键盘、鼠标、屏幕数据训练 AI,认为员工平均智力高于外包,可更快提升 Llama 的编码能力。作者指出 20 天后 8000 名员工收到裁员邮件,并批评这是把员工当免费高质量训练数据、用完就裁的做法。

    引用More Perfect Union (@MorePerfectUS)@MorePerfectUS

    LEAKED AUDIO: In an all-hands meeting on April 30, Mark Zuckerberg tells employees that he's training AI on them ahead of mass layoffs. "The AI models learn from watching really smart people do things... The average intelligence of the people who are at this company is significantly higher than the average set of people that you can get to do tasks. So if we're trying to teach the models coding, for example, then having people internally build tools or solve tasks that help teach the model how to code, we think is going to dramatically increase our model's coding ability faster than what others in the industry have the capability to do, who don't have thousands and thousands of extremely strong engineers at their company." Video

    推荐理由:引用泄露的全员会音频,呈现 Meta 用员工数据训练 AI 与随后裁员的关联叙事,可借此了解事件背景。

5月20日周三
  1. @kimmonismus69

    泄露的 Meta 4 月 30 日全员会录音显示,扎克伯格告诉员工,公司正在用他们训练 AI 模型,之后将进行大规模裁员。他称 Meta 工程师的平均智力显著高于外部可雇到的人,让内部人员构建工具或解决编码任务,能让模型的编码能力比缺乏大量顶尖工程师的竞争对手提升得更快。裁员预计在周三凌晨 4 点进行。

    引用More Perfect Union (@MorePerfectUS)@MorePerfectUS

    LEAKED AUDIO: In an all-hands meeting on April 30, Mark Zuckerberg tells employees that he's training AI on them ahead of mass layoffs. "The AI models learn from watching really smart people do things... The average intelligence of the people who are at this company is significantly higher than the average set of people that you can get to do tasks. So if we're trying to teach the models coding, for example, then having people internally build tools or solve tasks that help teach the model how to code, we think is going to dramatically increase our model's coding ability faster than what others in the industry have the capability to do, who don't have thousands and thousands of extremely strong engineers at their company." Video

    推荐理由:泄露录音呈现了 Meta 一边用工程师产出训练模型、一边推进裁员的内部逻辑,可作为观察大厂 AI 数据来源的个案。

  2. @swyx75

    Nick Joseph 发文欢迎 Andrej 加入预训练团队,Andrej 将组建一个专注用 Claude 加速预训练研究本身的团队。swyx 转发这条消息并评论 rsi is here。

    引用Nicholas Joseph (@nickevanjoseph)@nickevanjoseph

    Excited to welcome Andrej to the Pretraining team! He'll be building a team focused on using Claude to accelerate pretraining research itself. I can’t think of anyone better suited to do it — looking forward to what we build together!

    推荐理由:预训练团队迎来 Andrej 并将用 Claude 加速预训练研究本身,转发者把这条人事消息接到 RSI 的讨论上。