在我们的 Speech Agent Arena 中,人们进行盲测实时语音对话并选出自己更偏好的模型,Gemini 3.8 Live 首次亮相便位列第 2,Elo 为 1083,仅次于 Gemini 3.1 Flash Live(1096),领先于 GPT-Live-1(Sol, low)的 1053。它还成功完成了 93.2% 的任务,在任务成功率上排名第 2,仅次于 Grok Voice Think Fast 2.0 High 的 94.6%。Gemini 3.8 Live Extended Thinking (High) 的 Elo 为 990,任务成功率为 89.1%。
In our Speech Agent Arena, where people hold blind live voice conversations and pick the model they prefer, Gemini 3.8 Live debuts at #2 with an Elo of 1083, behind only Gemini 3.1 Flash Live (1096) and ahead of GPT-Live-1 (Sol, low) at 1053. It also completes 93.2% of tasks successfully, #2 on Task Success Rate behind Grok Voice Think Fast 2.0 High at 94.6%. Gemini 3.8 Live Extended Thinking (High) sits at Elo 990 with an 89.1% Task Success Rate.
来源:@ArtificialAnlys · x.com