跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-24AI 评分28
AI 导读

3/ 学习能否迁移到训练环境之外? 在 **Slay the Spire 2**——一个它未曾训练过的环境——中,dots3-note preview 必须理解不断变化的游戏状态,评估其选项,并在长序列决策中调整策略。 它不遵循固定计划,而是不断从每个结果中学习: **观察 → 评估 → 行动 → 学习 → 修正** 这就是长时程泛化的实际表现。

正文

3/ Can learning transfer beyond the training environment?

In **Slay the Spire 2**—an environment it was not trained on—dots3-note preview must understand a changing game state, evaluate its options, and adapt its strategy across a long sequence of decisions.

Instead of following a fixed plan, it keeps learning from each outcome:

**Observe → evaluate → act → learn → revise**

This is long-horizon generalization in action.

来源:@rohanpaul_ai · x.com