跳到正文
@omarsar0· @omarsar0 · X·· 22 天前AI 评分57
AI 导读

NVIDIA 论文比较多智能体系统中八种模型选择策略,在困难科学基准上覆盖路由、多数投票和 LLM-as-judge 设置。扩大不同开源模型池会提高理论最佳准确率,但实际准确率常低于池中单一最佳模型;使用同一模型多个副本效果更好,多数投票在最佳单模型上将 HLE 准确率从 29.4% 提升到 32.2%,而几乎所有混合模型组下降,从单一模型家族选择候选在八种策略中提升最大。

正文

Banger paper from NVIDIA.

It's on the topic of choosing which models go into a multi-agent system.

The team compared eight selection strategies, based on size, accuracy, answer diversity and error diversity, across routing, majority vote and LLM-as-judge setups on hard science benchmarks.

Larger pools of different open models raised the theoretical best-case accuracy. Achieved accuracy often fell below the single best model in the pool.

Using several copies of one model worked better.

Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined.

Choosing candidates from a single model family gave the largest improvement over a standalone model of all eight strategies.

Before adding another model to a router or ensemble, measure what it adds.

Paper: https://t.co/CzF2l8AOI7

Chat with Paper: https://t.co/ozsgU5cOZn

来源:@omarsar0 · x.com