跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-22AI 评分38
AI 导读

Google 提出 EnvHarness,让智能体训练环境随模型弱点自适应重塑,无需重建基准或验证器。EnvRigger 从 rollout 中找出弱点并生成 wrapper,仅在验证有效且可解时保留。在 SWE-bench Verified 上,同等 300 环境预算下智能体解决率 54.79%,高于原始环境的 52.13% 和生成环境的 50.37%。

正文

Another great Google paper.

Agent training has a ceiling: the agent improves, but the same environment stays frozen.

EnvHarness makes the environment adapt too, without rebuilding the benchmark or its verifier.

EnvHarness does this by reshaping an existing environment around whatever the agent is currently bad at, while keeping the original task and verifier intact.

If a coding agent skips tests, for example, the environment can reject submission until tests run; if it relies on a shortcut, that shortcut can be blocked.

EnvRigger finds these weaknesses from rollouts, writes the wrapper, then keeps it only if fresh runs show the new setup is both useful and solvable.

On SWE-bench Verified, with the same 300-environment budget, agents resolved 54.79% of issues, versus 52.13% with original environments and 50.37% with generated ones.

– arxiv. org/abs/2608.19880

Title: "EnvHarness: Awakening Static Worlds for Agent Learning"

来源:@rohanpaul_ai · x.com