AI 导读
3/ 学习能否迁移到训练环境之外? 在 **Slay the Spire 2**——一个它未曾训练过的环境——中,dots3-note preview 必须理解不断变化的游戏状态,评估其选项,并在长序列决策中调整策略。 它不遵循固定计划,而是不断从每个结果中学习: **观察 → 评估 → 行动 → 学习 → 修正** 这就是长时程泛化的实际表现。
正文
3/ Can learning transfer beyond the training environment?
In **Slay the Spire 2**—an environment it was not trained on—dots3-note preview must understand a changing game state, evaluate its options, and adapt its strategy across a long sequence of decisions.
Instead of following a fixed plan, it keeps learning from each outcome:
**Observe → evaluate → act → learn → revise**
This is long-horizon generalization in action.
来源:@rohanpaul_ai · x.com