一位 harness 构建者认为子智能体最适合并行化研究、代码审查和上下文管理,但协调问题是多子智能体架构崩溃的主因,成本在多数场景下不划算。他实测一个编排智能体加一个执行子智能体的组合效果最佳,增加更多子智能体会导致系统崩溃,且这一模式在混合模型家族时依然有效。他还指出,多智能体系统的进展有限,可能本质上是上下文工程问题——前沿模型在上下文过于多样时表现不佳。
If you build agent harnesses, this is important.
Should you avoid subagents, or can they be useful?
My thoughts as a harness builder:
I remember using subagents in Claude Code, and I mostly found them useful for parallelizing research. I didn't trust them for other things like coding.
In fact, I think parallelization, monitoring/tracking, and better context management are two of the best arguments for using subagents.
But are those reasons enough to justify the cost?
It depends. For code review, I think subagents are amazing and a great fit. And I like that you can do this efficiently with subagents.
Subagents work well if the orchestrator (manager) agent can coordinate the task and the subagents (executors) properly.
Like Eric, I have found that combining one orchestrator agent and one executor subagent typically works best right now. If you try, for instance, to add another subagent to the mix, things start to collapse.
What's been interesting is that these patterns work even when mixing model families. It feels like frontier models are trained to do this well.
Coordination is where multi-subagent architectures fall apart, and I think that's what Eric is pointing to. And the cost is just not worth it in most cases.
So when you see someone on X bragging bout their 100+, 2+ levels deep multi-agent system, you almost certainly know it's made up.
But it's surprised me that we haven't made much progress on subagents.
Although I have seen a few papers and engineering blogs sharing success using a form of message board or scratchpad with multi-agent systems. It's incredible how far harness engineering can take you. This tells me that maybe subagents could be a context engineering problem, i.e., frontier models don't do so well when context is too diverse.
This is interesting, as it might be that frontier models simply haven't been trained enough to be robust to this.
Which brings me to a point I have been raising more recently on avoiding using models to generate harnesses on the fly. They are cost-prohibitive and really hard to make them work on domain-specific tasks (see dynamic workflows from ant). But more on this another day.
I still think subagents are a useful primitive for agent harnesses. For long-horizon, complex tasks, I think they could be extremely useful for improving efficiency. For instance, subagents can explore experiments in parallel in research automation tasks.
For agent teams, I also think subagents remain relevant. But until we solve the cost or coordination problem, it will take time for the subagent pattern to be widely adopted.
I have more to share, but what has your experience been? Curious to know.
来源:@omarsar0 · x.com