ArcticSwarm 通过让部分智能体在搜索阶段互相隔离、之后再汇总评审,解决了多智能体过早达成共识的问题。在 BrowseComp-Plus 上搭配 Qwen 3.5-27B 达到 82.6% 准确率,去掉搜索隔离降至 78.8%,再去掉评审系统降至 74.5%。即便 40 次独立单智能体运行加多数投票也仅达 63.5%。
Letting research agents talk too early can make them follow the same wrong idea, so isolate some searches before review.
Multi-agent research works better when agents search independently before comparing notes, so delay collaboration until there is evidence to review.
The problem is once 1 agent finds a plausible answer, other agents can start searching around that same idea instead of testing different possibilities.
The paper calls this premature consensus.
ArcticSwarm fixes it by blocking selected agents from reading their peers while they search, then bringing the findings together for review.
On BrowseComp-Plus with Qwen 3.5-27B, it reached 82.6% accuracy.
Remove that search isolation and accuracy fell to 78.8%.
Remove the review system too, and it fell to 74.5%.
Even 40 independent single-agent runs with majority voting reached only 63.5%.
So the lesson is not "add more agents" or "make them communicate more."
For difficult research tasks without a reliable verifier, give agents room to explore different paths first, then challenge and verify the leading answer before the swarm agrees.
来源:@rohanpaul_ai · x.com