AI 导读
这很有启发性,但也引出了一个更深层的问题。 旧的 scaling 范式: 预训练阶段投入更多算力 -> 更聪明的模型 当前的 scaling 范式: 推理阶段投入更多算力 -> 更好的答案 下一个 scaling 范式: ? -> ? 如果小模型 + test-time scaling + 工具开始触及天花板,而未来又转回更大的模型、能在单次前向传播中解决更多问题,那下一个 scaling 轴究竟是什么?
正文
This is illuminating, but it also raises a deeper question.
The old scaling paradigm:
More compute during pretraining -> smarter model
The current scaling paradigm:
More compute during inference -> better answer
The next scaling paradigm:
? -> ?
If small models + test-time scaling + tools are beginning to hit a ceiling, and the future shifts back toward much larger models that can solve more problems in a single forward pass, what exactly is the next scaling axis?
来源:@dongxi_nlp · x.com