跳到正文
@omarsar0· @omarsar0 · X·· 13 天前AI 评分38
AI 导读

有意思的结果。这就是为什么我预计会有更多智能体工作负载跑在混合模型上。 来自 @TheUnbiasedCo 的 Pareto 26.9 会把请求发给多个前沿模型和开源模型,然后保留最佳答案。 在针对 30 个智能体任务的新评测中,Pareto 与 GPT-6 Astra 并列第一,而每个成功任务的成本约为后者的 1/3。 它完成任务的速度也快于 DeepSeek V4 Pro 和 GLM 5.3 Flash。

正文

Interesting results here. This is why I expect more agent workloads to run on blended models.

Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.

In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.

It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.

来源:@omarsar0 · x.com