Google 发布音频模型 Gemini 3.5 Live Translate,支持 70+ 语言的低延迟实时语音翻译,通过近乎实时地处理流式语音来构建低延迟音频体验。该模型支持单会话内多语言输入、自动检测语种并即时开始翻译,采用原生音频处理以保留说话人的语调、语速与音高,并能过滤环境噪声以在嘈杂环境中保持对话清晰。
Our latest audio model, Gemini 3.5 Live Translate, takes real-time speech translation to the next level for developers by delivering low-latency translation across 70+ languages.
By processing speech as it streams in near real time, the model enables devs to build low-latency audio experiences with:
— Multilingual input: Understands multiple languages in a single session without needing to adjust settings.
— Auto-detection: Identifies the spoken language and begins translation instantly.
— Native audio processing: Generates more natural-sounding speech that preserves speakers' intonation, pacing, and pitch.
— Noise robustness: Filters out ambient noise for clearer conversation in loud environments.
Video
来源:@googleaidevs · x.com