AI 导读
LiveKit 基于 Gemini 3.5 Live Translate 构建了多语言多人视频通话演示,每人选择自己的语言即可实时听到对方以自己语言说出的语音,代码已开源在 GitHub。被转发的 Google 开发者账号称该音频模型以近实时方式处理流式语音,支持 70+ 种语言、自动检测说话语言、保留说话者语调语速音高的原生音频处理,以及过滤环境噪声。
正文
We built a live multilingual, multi-person video call with Gemini 3.5 Live Translate on LiveKit. Everyone picks their language, speaks naturally, and hears each other in real time in their language of choice.
Watch the demo and check out the open source repo: github.com/livekit-examples/…
Video
Our latest audio model, Gemini 3.5 Live Translate, takes real-time speech translation to the next level for developers by delivering low-latency translation across 70+ languages. By processing speech as it streams in near real time, the model enables devs to build low-latency audio experiences with: — Multilingual input: Understands multiple languages in a single session without needing to adjust settings. — Auto-detection: Identifies the spoken language and begins translation instantly. — Native audio processing: Generates more natural-sounding speech that preserves speakers' intonation, pacing, and pitch. — Noise robustness: Filters out ambient noise for clearer conversation in loud environments. Video在 X 查看被引用的帖子
来源:@livekit · x.com