Salesforce AI 的论文发现,用更强专家的完整轨迹微调已与弱模型协同演进的智能体框架,会让弱模型在七项企业任务上表现下降 4 到 30 分,跨 Qwen3-Coder 和 Gemma 4 两个模型家族。
Nice paper from Salesforce on co-evolving harnesses and models.
Harness engineering is a hot topic right now. So this is a great read.
(bookmark it)
Salesforce evolved a harness with a weak model across seven enterprise agent tasks, then trained that model on a stronger expert's full trajectories under the same harness.
Performance dropped on all seven tasks, by 4 to 30 points across Qwen3-Coder and Gemma 4.
The same fine-tuning helps under the unevolved harness. So the harness is what changes the outcome.
Their analysis points at model-harness fit.
Imitation transfers knowledge and increases scaffold usage, but the weaker model adopts the expert's planning strategy without the competence to execute it, and it no longer matches a harness that was evolved around its own native planning style.
The fix is to stop copying whole trajectories.
A meta-level agent finds the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn. That keeps the model's planning style intact and combines the gains from harness evolution and weight updates.
Paper: https://t.co/gG3J6MnvT0
来源:@omarsar0 · x.com