AI 导读
传统的标量价值头用固定的前向传播来估计价值。 TEMPO 则使用智能体价值模型,可以在测试时扩展其推理和工具使用。 在 ARC-AGI-3 上,TEMPO 得分: → 比基线 checkpoint 高 31.5% → 比 GRPO 高 20.6%
正文
A traditional scalar value head estimates value with a fixed forward pass.
TEMPO instead uses an agentic value model that can scale its reasoning and tool use at test time.
On ARC-AGI-3, TEMPO scored:
→ 31.5% above the baseline checkpoint
→ 20.6% above GRPO https://t.co/Wvbb90qp0k
来源:@omarsar0 · x.com