AI 导读
Meta Superintelligence Labs 发布 Muse Voice Transcribe,称其为首个实时音频感知模型,支持实时流式 ASR、20+ 说话人分离与端点检测。该模型支持多语言并无缝切换语码,可通过语言、关键词和上下文偏置提升准确率。Meta 称其在 ArtificialAnlys 流式语音转文字及公开说话人分离基准上排名第一。
正文
Introducing Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs.
Muse Voice Transcribe delivers real-time streaming ASR, diarization with 20+ speakers, and endpointing. It’s multilingual with seamless code-switching and improves accuracy with language, keyword, and context biasing.
The model ranks first on @ArtificialAnlys streaming speech-to-text and on public diarization benchmarks.
来源:@AIatMeta · x.com