跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-22AI 评分42
AI 导读

ClawGym II 提出无需访问框架内部即可对智能体进行 RL 训练,将 Claude Code 或 OpenClaw 视为黑盒纳入训练循环。框架在沙箱中运行未修改的 OpenClaw 或 Claude Code,在服务边界拦截模型调用,并将碎片化调用重建为前缀树轨迹供 PPO 或 GRPO 优化。

正文

You can now RL-train an agent through the same complex harness it will actually run in, without needing access to the harness internals.

ClawGym II shows that Claude Code or OpenClaw can be treated as a black box and still become part of the RL training loop.

The framework runs OpenClaw or Claude Code unchanged inside sandboxes, intercepts model calls at the serving boundary, and rebuilds fragmented calls into prefix-tree trajectories that PPO or GRPO can optimize.

That lets the model learn through the harness without the training stack reproducing its tool routing, retries, context management, or subagents.

With Qwen3-30A3B, this raised ClawGym-Bench Pass@1 by 9.98 points through OpenClaw and 14.81 points through Claude Code.

Mix-harness training also worked: a policy trained from OpenClaw and Claude Code matched or slightly beat the corresponding single-harness models under both execution systems.

The paper also reports gains on JobBench and OfficeQA, so the setup extends beyond ClawGym-style tasks.

– arxiv. org/abs/2608.16798

Title: "ClawGym II: Exploring Black-Box RL on Agent Harness"

来源:@rohanpaul_ai · x.com