Gemini 3.8 Live 在 Big Bench Audio 基准上的平均首音频时间(Time to First Audio)为 1.18 秒,Extended Thinking (High) 为 1.35 秒——两者都明显领先于 Gemini 3.1 Flash Live High 的 2.99 秒。这使标准模型略微领先于 GPT-Realtime-2.1 High(1.21 秒)、GPT-Live-1 (Sol, low)(1.24 秒)和 GPT-Live-1 (Astra, medium)(1.34 秒),而 Grok Voice Think Fast 2.0 High(0.70 秒)是我们 Speech to Speech Index 前 5 名中唯一低于 1 秒的模型。
Gemini 3.8 Live's average Time to First Audio on the Big Bench Audio benchmark is 1.18 seconds, and 1.35 seconds for Extended Thinking (High) - both well ahead of Gemini 3.1 Flash Live High at 2.99 seconds. That puts the standard model marginally ahead of GPT-Realtime-2.1 High (1.21s), GPT-Live-1 (Sol, low) (1.24s) and GPT-Live-1 (Astra, medium) (1.34s), with Grok Voice Think Fast 2.0 High (0.70s) the only model in the top 5 of our Speech to Speech Index under 1 second.
来源:@ArtificialAnlys · x.com