AWS AI Labs 新论文测量了智能体运行中切换模型的实际代价,作者称之为"交接税"。在 Claude 与 GPT 模型配对测试中,全轨迹升级到强模型只能挽回弱强模型间不到一半的质量差距,还附带显著成本溢价;而降级切换反而能落在成本与质量更优的平衡点上。削减弱模型轨迹信息可提升升级质量,移除强模型轨迹则会损害降级质量。
Great new paper from AWS on agent handoff tax.
If you build agents today, you need to understand the so-called handoff tax.
(bookmark it)
Escalating to a stronger model mid-run is usually the resort when a cheap agent stalls.
New work from AWS AI Labs measures how much that switch actually costs.
Coding agents run for dozens of model calls, so teams escalate when a weak model struggles and downshift once the hard reasoning is done. Every switch forces the receiving model to continue a trajectory another model wrote.
Across pairs of Claude and GPT models, full-trajectory escalation recovers less than half the quality gap between the weak and strong model while adding a substantial cost premium. The authors call that penalty the handoff tax. Downshifting lands at a much better cost-quality point.
Cutting the weak model's trajectory information improves escalation quality, while removing the strong model's trajectory hurts downshift quality.
Paper: https://t.co/59ozKgugB8
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com