一项模型蒸馏研究发现,蒸馏出的学生模型复制的是教师的推理方式而非知识本身,因此教师血统远比教师基准分数重要。用学生自身基座模型构建的较弱教师,比来自不同家族的更强教师效果更好;训练数据几乎不影响结果,用小学水平数学题即可获得超 80% 的收益。教师甚至能教会学生自己解不出的问题,因为传递的是推理习惯。
Interesting study on model distillation.
A distilled student model copies how its teacher reasons rather than what it knows, so teacher lineage matters far more than teacher benchmark scores.
A weaker teacher model built from your student's own base model will take it further than a stronger teacher from a different family.
This paper finds, the data barely matters. Train on problems the teacher always solves, never solves, or a random mix, and you land in the same place. Even grade-school math gets you over 80% of the benefit.
A teacher can teach a problem it can't solve itself, because what's moving across is a habit of reasoning.
来源:@rohanpaul_ai · x.com