AI 导读
2/ 新环境中的测试时学习 当 dots3-note preview 进入一个陌生环境时,它会探索、形成假设、验证假设,并在证据证明其错误时更新自己的理解。 它还能改写自己的记忆——记录有用的发现,并将修正带入后续决策。 **探索 → 假设 → 验证 → 更新记忆 → 再试一次** 通过 RL,该模型不仅学会了如何行动,还学会了什么值得记住。
正文
2/ Test-Time Learning in Novel Environments
When dots3-note preview enters an unfamiliar environment, it explores, forms hypotheses, tests them, and updates its understanding when the evidence proves it wrong.
It can also rewrite its own memory—recording useful discoveries and carrying corrections into later decisions.
**Explore → hypothesize → test → update memory → try again**
Through RL, the model learns not only how to act, but also what is worth remembering.
来源:@rohanpaul_ai · x.com