跳到正文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分57
AI 导读

Evolvent AI 发布开源研究环境 RSIGym,并推出 RSI-Index 衡量前沿智能体同时改进目标模型权重与 harness 的能力。

正文

Claude Opus 5 nearly tripled a Qwen model's SWE-bench Verified score while working as an automated AI researcher.

@Evolvent_AI 's newly released open-source RSIGym made that measurement possible by letting an AI agent retrain a model and rewrite its harness inside one budgeted environment.

shows that frontier AI agents can substantially improve another AI model when training, serving, and testing come as ready-made services.

Agents work from CPU-only containers and call remote services for LoRA fine-tuning, model serving, benchmarking and sandboxes, all charged against a per-run budget.

In the main test, 6 frontier agents started from Qwen3.5-35B-A3B-Base and a minimal harness, with $500 of services per benchmark.

If you're improving an agent, read its failure logs and fix the harness first: heavier training scored lower in 8 of 10 small tests.

引用Fanqing@FanqingMengAI
💡Our view: RSI is a systems engineering problem, not just a model problem. Progress depends on the environment a research agent works in: what resources it can call, what it can change, and how it runs experiments. That environment should reflect real production workflows and be easy for agents to use. At @Evolvent_AI, we built RSIGym around Everything as a Service. Agents can call established research services and iterate on data, training settings, and harness code. We also introduce RSI-Index to measure how well frontier agents jointly improve a target model’s weights and harness. Across 6 research agents and 5 benchmarks, Opus 5 leads at 0.4809. 🧵 We’re releasing the code, experiment configurations, research trajectories, and evaluation logs. Try your own improvement methods with RSIGym. 🌐 Website: https://rsi-index.ai/ 💻 GitHub: https://github.com/evolvent-ai/RSIGym 📄 arXiv: https://arxiv.org/abs/2610.10310
在 X 查看被引用的帖子

来源:Rohan Paul · x.com