跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-25AI 评分42
AI 导读

给语言模型加新模态只需训练编码器与主干之间的投影层,无需改动主干权重,已上线能力不会退化。3D 测试中冻结主干模型匹配甚至超越联合微调版本,训练速度还快一倍;而联合微调的 Llama 主干在 GSM8K 上从 86.96% 崩到 0.61%。论文标题为 "Projector Is All You Train"。

正文

You can teach a language model a new modality without ever touching the backbone's weights.

Train only the projector between encoder and backbone, and the capabilities you already shipped cannot regress.

To add a new modality to a language model, you only need to train the small projector that maps the encoder's output into the model's embedding space — fine-tuning the language model itself adds nothing reliable and wrecks its existing skills.

In their 3D tests, frozen-backbone models matched or beat jointly fine-tuned ones while training twice as fast, and the jointly fine-tuned Llama backbone collapsed from 86.96% to 0.61% on GSM8K.

– arxiv. org/abs/2608.19726

Title: "Projector Is All You Train"

来源:@rohanpaul_ai · x.com