跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 17 天前AI 评分46
AI 导读

FrogNano 从 Qwen3.5-4B 出发,用 RL 在约 1500 个合成软件工程任务上训练,在 SWE-bench Verified 上达到 61.5%,且未使用前沿模型蒸馏。其关键是在线任务合成:随模型能力提升持续生成有挑战但可学习的新题,让课程与智能体共同进化。改用更简单的 5 工具接口后,基座模型从 8.3% 提升到 37.2%,5 轮训练后达到 61.5%。

正文

A 4B coding agent reached 61.5% on SWE-bench Verified without frontier-model distillation by combining a simpler tool interface with synthetic tasks that keep adapting to what the model can currently learn.

FrogNano starts from Qwen3.5-4B and is trained with RL on about 1,500 synthetic software-engineering tasks.

The important part is how those tasks are chosen.

As the model improves, the system generates fresh problems that are challenging but still learnable, so the curriculum improves with the agent.

The interface matters just as much.

Switching to a simpler 5-tool setup moved the base model from 8.3% to 37.2% on SWE-bench Verified.

After 5 rounds, FrogNano reached 61.5%.

– arxiv. org/abs/2609.07925

Title: "FrogNano: Training a 4B Coding Agent via Online Task Synthesis"

来源:@rohanpaul_ai · x.com