AI 导读
有意思的结果。这就是为什么我预计会有更多智能体工作负载跑在混合模型上。 来自 @TheUnbiasedCo 的 Pareto 26.9 会把请求发给多个前沿模型和开源模型,然后保留最佳答案。 在针对 30 个智能体任务的新评测中,Pareto 与 GPT-6 Astra 并列第一,而每个成功任务的成本约为后者的 1/3。 它完成任务的速度也快于 DeepSeek V4 Pro 和 GLM 5.3 Flash。
正文
Interesting results here. This is why I expect more agent workloads to run on blended models.
Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.
In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.
It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.
来源:@omarsar0 · x.com