斯坦福大学论文提出 CLEAR,通过按提示词动态路由安全模块来缓解安全对齐带来的能力损失。该方法冻结原始模型,用一个小型门控网络读取每条提示词,决定开启多少独立安全模块;良性提示词几乎不启用安全模块,直接由未改动的原模型处理。
Safety alignment usually costs utility because the aligned weights apply to every prompt;
New Stanford univ paper finds that scaling the safety update per prompt recovers much of what global tuning gives up.
The problem is that a safety fine-tune changes the model for every input. Harmful or not, every prompt now runs through a safer but weaker model.
CLEAR, proposed in this paper, leaves the original model frozen. A small gate reads each incoming prompt and decides how much of a separate safety module to switch on.
Benign prompts get almost none of it, so they run on the untouched model.
– arxiv. org/abs/2608.21278
Title: "CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment"
来源:@rohanpaul_ai · x.com