Meta 的论文对比字节级模型与 token 模型在蒸馏和算力缩放下的表现,发现字节模型起步落后但随算力增长反超,拟合缩放律预测 End-Of-Token 模型最终比蒸馏 token 模型领先最多 4%。
Banger paper from Meta.
This work shows that byte-level models start out behind token models and then pass them as compute grows.
They show this for distilled 1B models trained on up to 1 trillion bytes.
To distill a byte student from a token teacher, they convert the teacher's token logits into byte logits, either approximately (Marginalize-It) or exactly (End-Of-Token).
Token models lead at low compute but plateau.
Byte models reach a higher ceiling, and the fitted scaling laws predict the End-Of-Token model ends up to 4% ahead of the distilled token model.
The byte models also match the distilled token model with one-sixth of the training data, and a 256-entry vocabulary cuts teacher-logit storage to about a fifth.
Paper: https://t.co/NBW7r33GMH
Chat with Paper: https://t.co/ibuIi7to7g
来源:@omarsar0 · x.com