跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-20AI 评分35
AI 导读

传统的标量价值头用固定的前向传播来估计价值。 TEMPO 则使用智能体价值模型,可以在测试时扩展其推理和工具使用。 在 ARC-AGI-3 上,TEMPO 得分: → 比基线 checkpoint 高 31.5% → 比 GRPO 高 20.6%

正文

A traditional scalar value head estimates value with a fixed forward pass.

TEMPO instead uses an agentic value model that can scale its reasoning and tool use at test time.

On ARC-AGI-3, TEMPO scored:

→ 31.5% above the baseline checkpoint
→ 20.6% above GRPO https://t.co/Wvbb90qp0k

来源:@omarsar0 · x.com