跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-30AI 评分37
AI 导读

斯坦福大学论文提出 CLEAR,通过按提示词动态路由安全模块来缓解安全对齐带来的能力损失。该方法冻结原始模型,用一个小型门控网络读取每条提示词,决定开启多少独立安全模块;良性提示词几乎不启用安全模块,直接由未改动的原模型处理。

正文

Safety alignment usually costs utility because the aligned weights apply to every prompt;

New Stanford univ paper finds that scaling the safety update per prompt recovers much of what global tuning gives up.

The problem is that a safety fine-tune changes the model for every input. Harmful or not, every prompt now runs through a safer but weaker model.

CLEAR, proposed in this paper, leaves the original model frozen. A small gate reads each incoming prompt and decides how much of a separate safety module to switch on.

Benign prompts get almost none of it, so they run on the untouched model.

– arxiv. org/abs/2608.21278

Title: "CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment"

来源:@rohanpaul_ai · x.com