跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-02AI 评分58
AI 导读

Meta 发布 Muse Voice Transcribe 实时语音转写模型,最终转写词错误率最低为 3.1%,并支持自适应延迟。该模型在同一个流式模型内原生完成说话人分离与端点判断,能学习何时等待、何时输出词以及一轮对话何时结束。音频按 80ms 分块处理,每块之后决定输出文本还是继续监听,目前已通过 Meta Model API、Mac 版 Meta AI 和 Muse Code 提供。

正文

Love this, another huge release from Meta.

Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest 3.1% final-transcription word error rate with adaptive delay.
which is significantly ahead of competing models.

- The big deal is that Muse does something beyond just transcribing speech; it learns when to wait, when to commit a word, when a speaker changes, and when a turn is actually over, all inside the same streaming model. That makes it much closer to a real-time perception layer for voice agents than a conventional speech-to-text API.

- Muse processes audio in 80ms chunks and chooses after each chunk whether to emit text or keep listening.

Meta made it available through Meta Model API, Meta AI for Mac, and Muse Code

引用@finkd@finkd
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model. https://t.co/LViMDSkbim
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com