跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-21AI 评分44
AI 导读

Google 提出 EnvHarness 与 EnvRigger,用可编程插件层重塑静态智能体训练环境,且每个重塑环境保留原始 verifier 以保证训练安全。EnvRigger 将策略视为黑盒,读取执行轨迹、诊断缺陷并合成 harness 组件,再用新 rollout 验证。在四个领域的五个 benchmark 上,held-out 实例最高提升 9.0 分,执行步数减少 9.8%。

正文

Impressive research from Google on building better environments for agents.

Training environments for agents are hand-built and go stale. The agent improves, the environment does not, and it's not able to see the agent's weaknesses in the first place.

EnvHarness wraps a static environment in a programmable plug-in layer that reshapes its behavior without touching the underlying logic. Every reshaped environment keeps its original verifier; this is what makes the reshaping safe to train on.

EnvRigger treats the policy as a black box, reads its execution trajectories, synthesizes harness components aimed at the diagnosed flaws, then validates them with fresh rollouts.

Across five benchmarks in four domains, up to 9.0 points better on held-out instances with 9.8% fewer execution steps.

Paper: https://t.co/sCeXq0jdMg

Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX

来源:@omarsar0 · x.com