跳到正文

#数据/训练

今日 8 条
今天10月2日周五
  1. Thomas Wolf53

    Thomas Wolf 发推调侃 Karpathy 从 X 消失后,Ben Affleck 开始讲微调方法,称要先冻结基座权重、学习率用 2e-4。引用内容介绍 Affleck 通过解冻权重、只训练最后的电影级层来微调开放视频模型,其创办的 InterPositive 自建 8 个月数据集,并被 Netflix 以 5.87 亿美元现金收购。

    引用Rohan Paul@rohanpaul_ai

    Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards. for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash. He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training. Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for. ---- From "Bloomberg Live" YouTube channel, (link in comment)

  2. Hugging Face Blog43

    AutoSynthData:为 Enterprise Agents 生成训练数据

    ServiceNow CoreAI 构建了 AutoSynthData,利用目标模型的失败案例和更强教师的成功轨迹,自动生成并验证新的可执行训练任务,形成随模型能力动态调整的课程。该方法在 EnterpriseOps Gym 环境中验证,通过能力规格卡生成多样化任务,避免直接使用原始提示词和轨迹,并已发布相关数据集。

10月1日周四
  1. Google Cloud: Databases43

    Google Cloud Managed Lustre 推出 6 美分/GB*月 Dynamic Tier,降低高性能并行文件系统门槛

    Google Cloud Managed Lustre 新增 Dynamic Tier,以 6 美分/GB*月提供热数据亚毫秒延迟,容量可线性扩展至 80 PB,且不单独收取介质、数据移动和元数据 IOPS 费用。该层支持多轮训练、写入密集检查点、快速检查点恢复和交互式开发,平均读延迟约 300µs,比替代分布式文件系统响应快至 4 倍。

9月30日周三
  1. elsewhere articles34

    心资本韩彦 SuperReturn Asia 演讲:2026至2028年中国硬科技早期投资窗口在哪

    心资本创始合伙人韩彦在 SuperReturn Asia 2026 演讲中判断,2026至2028年是中国硬科技早期投资窗口:IPO市场回暖、退出路径变清晰,但愿意支持早期企业的资本尚未充分回流。他援引数据称,2026年上半年香港85宗上市募资约2,100亿港元、同比增长92%,中国VC/PE投资金额4,310亿元人民币、同比增长173%,其中44.7%投向AI。

  2. elsewhere articles67

    数据公司估值飙升背后:给模型做练习题的生意逻辑

    作者观察到一批成立不满一年的数据公司估值快速增长,如 UniPat 达 25 亿美金,认为数据成为模厂竞争的核心原料。文章把数商业务概括为给模型制作练习题,即采集专家解题轨迹或搭建 RL environment,并分析其同时是信任市场、信息市场和油水很足的市场,还提及水下蒸馏与爬取数据的灰色操作,以及字节、阿里、腾讯等买方梯队。

    推荐理由:作者以一手行业观察拆解数据公司高估值背后的生意本质,包括信任、信息和关系三个维度,读来能理解这条赛道的运转逻辑。

  3. Hugging Face Daily Papers35

    邻近监督更优:Neighborhood OPSD 自蒸馏方法提升数学推理模型性能

    Neighborhood OPSD(N-OPSD)通过局部参数扰动构建冻结专家池,将参考对齐修正转化为学生可用的监督信号,在 AIME 2024、AIME 2025 和 HMMT February 2025 三项基准上,将 Average@12 较标准 OPSD 分别提升 2.75、1.67 和 1.94 分(对应 Qwen3-1.7B、4B、8B)。推理时仅使用蒸馏后的学生模型。

  4. Hugging Face Daily Papers46

    合成预预训练在规模扩展下依然有效,但并非语法先验

    一项覆盖 500M 至 7B 参数、PT 预算最高 100B token 的研究显示,合成数据预预训练(PPT)的 token 效率增益在规模扩展后依然存在,3B 规模下可节省至少 21B PT token。但研究未发现一致证据表明这些增益来自语法先验,下游表现与语法可接受性在各模型规模上并不一致。增益实际来自提升长程检索能力的 PPT 任务,且对 PT 数据混合方式稳健,仅在缺少网页文本时减弱。

  5. Hugging Face Daily Papers36

    DARA:面向多奖励强化学习的密度感知奖励聚合方法

    针对多奖励强化学习中各目标学习进度不均的问题,研究者提出密度感知奖励聚合方法 DARA,通过逆平方根密度校正,为激活频率较低的奖励信号赋予更高权重。在工具调用任务上,DARA 最多减少 26% 训练步数达到高格式合规率;在数学推理任务上,最多减少 65% 步数接近长度合规饱和,最终性能与 GDPO 相当。

  6. Hugging Face Daily Papers37

    AutoDataBench:面向加速自动研究的数据中心测试平台

    AutoDataBench 是一个隔离数据因素、系统评估 LLM「数据智能」的受控测试平台,覆盖数据诊断、组织与构建三类优化任务,在固定非数据因素下考察前沿 LLM 通过迭代实验改进训练数据的能力。研究还对比训练前预测与实际结果,检验 LLM 是否具备数据效应推理能力,并发现复用 AutoDataBench 轨迹进行 mid-training 可提升下游编码性能。

9月29日周二
  1. Thomas Wolf56

    modded-nanogpt 传入新的历史纪录 39.9 秒,较此前 67.6 秒快 27.7 秒,核心思路是在单个 flop 级别做稀疏优化而非只优化矩阵乘法。主要手段包括采样 softmax(约 8 秒)、稀疏 n-gram 嵌入更新与优化器状态、稀疏通信、最后 300 步 EMA(约 4 秒)、新优化器 Anvil2(约 1 秒)等,稀疏嵌入参数扩展到 65B,占本次提升的 25%。详见 https://github.com/KellerJordan/modded-nanogpt/pull/360 和 https://hyperstition.cc/training-nanogpt-in-39-9-seconds。

    引用Larry Dial@classiclarryd

    New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds

9月28日周一
9月27日周日
9月24日周四
  1. inclusionAI Hugging Face models60

    inclusionAI 开源 Ling-mini-2.0:16B 总参数 MoE 模型,激活仅 1.4B

    inclusionAI 开源 Ling 2.0 系列首个模型 Ling-mini-2.0,总参数 16.26B、每 token 激活 1.4B(非嵌入 789M),采用 1/32 激活比 MoE 架构,称可达到 7–8B dense 模型的等效性能,在 H20 上简单 QA 场景生成速度超过 300 token/s,支持 128K 上下文(YaRN)。

    推荐理由:官方发布了完整的参数配置、推理速度、训练吞吐和预训练检查点,读者可以据此评估小激活 MoE 在端侧和继续训练中的可用性。

  2. inclusionAI Hugging Face models60

    inclusionAI 发布万亿参数推理模型 Ring-2.6-1T

    inclusionAI 发布万亿参数旗舰推理模型 Ring-2.6-1T,主打真实生产环境中的 Agent 执行与复杂推理,上下文长度由 128K 扩展到 256K(YaRN),采用 MIT License 开源。

    推荐理由:官方发布页给出 Agent 执行、推理档位和异步 RL 训练的具体做法与基准数字,读者可据此评估其在生产场景的适配价值。

9月23日周三
  1. ByteByteGo40

    如何定制模型让它学会新技能:从提示词、RAG 到 LoRA 与 QLoRA 微调

    当提示词和 RAG 无法消除模型在特定任务上的反复偏差时,就需要通过微调把期望行为固化进模型参数。文章梳理了定制模型的路径:先用 few-shot prompting 和检索增强生成(RAG)补充信息,再用监督微调(SFT)训练模型,而 LoRA 和 QLoRA 通过减少需要调整的参数和显存占用,让微调现有模型更易落地。

  2. Anthropic Newsroom76

    Anthropic:Claude 发现类似 CRISPR 的新型酶系统 ART

    Anthropic 成立生命科学研究组和实验室,宣布 Claude 智能体在约 950 个智能体、21 小时、2.1 亿 token 的搜索后,自主发现一种与 DNA 重复序列相关的新型酶系统,命名为 array-associated reverse transcriptases(ART)。

    推荐理由:原文给出 Claude 智能体自主发现新酶系统的过程细节和预印本,读者可以据此了解 AI 驱动生物学研究的实际工作方式。

9月22日周二
  1. karminski-牙医39

    Qwen4 家族首次曝光,包含 Qwen4-Max、Qwen4-Flash & Qwen4-Plus 以及 Qwen4-27B,未来 Qwen 会训 5-10T 的模型。卧槽5-10T???

    引用Max For AI@MaxForAI

    🚨Qwen4家族首次曝光!! 刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族! 包含Qwen4-Max Qwen4-Flash&Qwen4-Plus 还有Qwen4-27B!!! 未来Qwen会训5-10T的模型

9月21日周一
  1. elsewhere articles26

    百曜科技联合《麻省理工科技评论》发布 AIVC 深度研究报告

    百曜科技联合《麻省理工科技评论》中国团队发布《AI 虚拟细胞(AIVC)技术趋势、产业生态与应用前景研究报告》,将 AIVC 定义为 AI 时代生命科学的新型基础设施。报告指出产业竞争正从模型规模转向"数据—模型—实验"干湿闭环,并梳理了泛化能力、私有数据与干湿闭环三大趋势。百曜科技今年 6 月已发布基于 LLM-JEPA 架构的 AI 虚拟细胞世界模型 AURA CellOS。