AI 导读
模型在不同价位上都能实现出色的对话偏好和任务成功率。Gemini 3.1 Flash Live Preview - Minimal 以 1,046 Elo 领跑整体偏好,输入音频成本为每小时 $1.50,任务成功率为 74.6%;而 Grok Voice Think Fast 2.0 High 以 94.7% 领跑任务成功率,每小时 $4.80。GPT-Realtime-2.1 High 以 91.5% 的任务成功率紧随其后,每小时成本为 $10.75。
正文
Models achieve strong conversational preference and task success at materially different price points. Gemini 3.1 Flash Live Preview - Minimal leads overall preference at 1,046 Elo with a cost of $1.50 per hour of input audio and a 74.6% Task Success Rate, while Grok Voice Think Fast 2.0 High leads Task Success Rate at 94.7% at $4.80 per hour. GPT-Realtime-2.1 High follows on task success at 91.5% and costs $10.75 per hour.
来源:@ArtificialAnlys · x.com