跳到正文
原文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-27精选AI 评分69
AI 导读

Google 发布 Gemini 3.5 Transcribe,一款不再逐字记录、而是输出用户本意的语音转写模型,现已通过 API 开放。

推荐理由

原文列出两个 API 端点各自的能力与时长限制,读者可据此判断语音转写在工作流中的可用边界。

正文

Incrdible release from Google, and one more friction gone from ditching my keyboard.

Also looks like, Gemini 3.5 Transcribe will put real pressure on many startups producing clean, well-formatted voice-to-text.

you speak into Gemini 3.5 Transcribe, and instead of directly writing down every word literally, it will produce the text you actually intended.

Its smart mode resolves false starts, strips filler, formats numbers and dates, and biases recognition toward custom vocabulary instead of returning speech word-for-word.

It is available through API right now. There are 2 endpoints:

- gemini-3.5-transcribe: recorded audio, up to 1 hour, with timestamps, speaker identification, custom vocabulary and smart transcription.

- gemini-3.5-transcribe-live: real-time WebSocket streaming, with smart transcription and sub-second interaction latency; individual live sessions currently max out at 10 minutes.

引用@sundarpichai@sundarpichai
Say hello to Gemini 3.5 Transcribe! - Build apps that understand user speech / intent, even w/ multiple speakers! - Auto-detection of 85+ languages out of the box - Custom vocab adaptation for specialized jargon... SGTM:) API available now in @GoogleAIStudio and Gemini Enterprise, or try it in the Gemini app on macOS or Rambler on Android! More details: https://t.co/AduutCb3M7
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com