MoE 强化学习中的专家空间探索(ESRL)
Expert-Space Exploration in MoE Reinforcement Learning
阅读原文
本站未展示全文,请前往来源网站阅读。
AI 导读
研究者提出 Expert-Space Exploration Reinforcement Learning(ESRL),一个显式探索 MoE 模型专家路由空间的架构感知框架,通过保留高置信专家作为锚点、将随机路由限制在候选池内并按 router 熵自适应调整扰动强度来提升 rollout 多样性。
来源:Hugging Face Daily Papers · arxiv.org