自然语言自编码器为 LLM 激活生成解释
Natural Language Autoencoders Produce Explanations of LLM Activations
阅读原文
本站未展示全文,请前往来源网站阅读。
AI 导读
自然语言自编码器(Natural Language Autoencoders)可用于生成对 LLM 激活的解释。该方向尝试把模型内部激活转化为自然语言描述,以提升可解释性。
来源:transformer-circuits.pub(经 Hacker News) · transformer-circuits.pub