跳到正文
@omarsar0· @omarsar0 · X·· 24 天前AI 评分51
AI 导读

Meta 的论文对比字节级模型与 token 模型在蒸馏和算力缩放下的表现,发现字节模型起步落后但随算力增长反超,拟合缩放律预测 End-Of-Token 模型最终比蒸馏 token 模型领先最多 4%。

正文

Banger paper from Meta.

This work shows that byte-level models start out behind token models and then pass them as compute grows.

They show this for distilled 1B models trained on up to 1 trillion bytes.

To distill a byte student from a token teacher, they convert the teacher's token logits into byte logits, either approximately (Marginalize-It) or exactly (End-Of-Token).

Token models lead at low compute but plateau.

Byte models reach a higher ceiling, and the fitted scaling laws predict the End-Of-Token model ends up to 4% ahead of the distilled token model.

The byte models also match the distilled token model with one-sixth of the training data, and a 256-entry vocabulary cuts teacher-logit storage to about a fifth.

Paper: https://t.co/NBW7r33GMH

Chat with Paper: https://t.co/ibuIi7to7g

来源:@omarsar0 · x.com