跳到正文

#数据/训练

今日 0 条
9月30日周三
9月24日周四
9月23日周三
9月12日周六
  1. Dwarkesh Patel56

    Dwarkesh 对谈 John Schulman、Beren Millidge 与 Charlie O'Neill:递归自我改进还有多远

    Dwarkesh Patel 与 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman、Baseten 模型训练负责人 Charlie O'Neill 长篇对谈,逐段讨论递归自我改进(RSI)最可能失败的技术原因、中国实验室的追赶路径、自动化 AI 研究者的训练方式以及长时程 RL 能否带来 AGI。

8月31日周一
8月24日周一
  1. Andrew Ng50

    在捍卫 AI 开放性的斗争中,Marin 项目是模型训练开放性的一份珍贵示范——开放代码、数据、配方,甚至实验结果。公开分享 AI 研究曾是常态;我很感激 @percyliang 的开放实验室做法。

    引用Percy Liang@percyliang

    🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.

8月21日周五
  1. jietang44

    精彩评论:FLOPs 是智能;参数是知识!

    引用Liam Fedus@LiamFedus

    An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!

8月19日周三
8月17日周一
8月12日周三
8月8日周六
  1. Dwarkesh Patel58

    Dwarkesh Patel 提出持续学习时代的 8 项 AI 预测

    Dwarkesh Patel 提出持续学习到来后的 8 项预测,认为模型仅在会话间写 Markdown 无法积累执行整份工作所需的经验,经验必须沉淀进权重。他据此推论部署前安全检查将失效、对齐技术需重构、领先实验室将靠部署数据加速拉开差距,并以 Anthropic 内部自 2 月起使用 Mythos、6 月才公开发布的 4 个月差距为例说明先发部署的学习优势。

7月7日周二
6月26日周五
  1. Dwarkesh Patel65

    Dwarkesh Patel:下一个突破是 AI 在工作中学习

    Dwarkesh Patel 撰文认为,实验室押注的 RLVR 训练未必能泛化到无法在数据中心内复现的现实领域,因为训练样本效率低且缺少可重放模拟器。

    推荐理由:作者论证 RLVR 难以泛化到不可复现的现实领域,并梳理 OPSD 与模拟训练等让模型在工作中持续学习的可能路径。

6月20日周六
  1. Dwarkesh Patel59

    Dwarkesh:AI 进步的核心是数据黑洞而非样本效率提升

    Dwarkesh Patel 撰文认为,近几年 AI 进步主要来自更宽更好的数据分布和 RL 合成数据,而非样本效率提升。他估算人类从出生到成年只接触约 2 亿 token,前沿模型却训练于数十至数百万亿 token,差距近百万倍;按 Chinchilla 常数,即使无限增加参数也只能将所需数据降低约 10 倍,说明人类处于不同的 scaling 曲线。

6月1日周一
5月18日周一
7月17日周四