ClawGym II 提出通过 OpenClaw 和 Claude Code 作为黑盒进行 RL 训练智能体的方法:在模型边界部署服务代理捕获 harness 的每次调用,再组织成前缀树,让 PPO 和 GRPO 能在恢复的多轮结构上优化。
Really interesting paper.
I recommend it to anyone interested in training agents using existing harnesses.
(bookmark it)
ClawGym II runs RL through OpenClaw and Claude Code as opaque boxes. A serving proxy sits at the model boundary and captures every call the harness makes, then those calls get organized into prefix trees so PPO and GRPO can optimize over the recovered multi-turn structure.
Qwen3-30A3B gains 9.98 points of Pass@1 through OpenClaw and 14.81 through Claude Code, stable across 200 to 400 optimization steps.
Mix-harness training pushes further. One model gets optimized jointly by heterogeneous harnesses, which points at policies that generalize across execution systems instead of overfitting to a single one.
Paper: https://t.co/Am5kH7TguU
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com