跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 21 天前精选AI 评分66
AI 导读

Google 发布 Gemini 3.8 Live 与 Gemini 3.8 Live Extended Thinking 两款实时多模态模型,称其拿下语音到语音综合评分第一。

推荐理由

材料给出 Gemini 3.8 Live 与竞品的每小时成本和 τ-Voice 得分对比,可用于判断语音智能体的性价比位置。

正文

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

And now it takes the #1 overall speech-to-speech score, with also a narrow quality lead but a much larger price advantage over its nearest frontier competitors (OpenAI's GPT-Live-1 (Astra, medium).

- Gemini 3.8 Live reaches genuinely frontier-level voice-agent performance at $0.84/hour, including a higher composite score than GPT-Realtime-2 High at roughly 80% lower measured cost.

- Both these models are multimodal live models, i.e. voice is the primary conversational interface, but visual input can provide additional context during the conversation. So both the models also process visual context and automatically switch among 97 supported languages during conversation.

- Extended Thinking scores 68.6% on τ-Voice, versus 67.9% for GPT-Live-1 Astra medium.

τ-Voice tests measures whether a voice model can actually finish a real multi-step task, not just sound natural or answer spoken questions. On this bench, the model has to hold a conversation, follow domain policies, use tools correctly, and reach the right outcome across airline, retail, and telecom customer-service scenarios.

引用@GoogleDeepMind@GoogleDeepMind
We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI. The models talk, think, and handle tasks in the background without breaking your flow. 🧵
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com