AI 导读
Alexandr Wang 宣布推出 muse voice transcribe,称这是其首个实时音频感知模型,在流式语音转写上达到 SOTA。该单一模型原生处理说话人分离与端点检测。随文图表显示,其最终转写错词率为 3.1%,低于图中列出的其他流式转写方案。
正文
1/ today we're rolling out muse voice transcribe, our first real-time audio perception model - SOTA in streaming speech-to-text. also handles speaker diarization and endpointing natively in a single model. https://t.co/fWx7uPxPn4
来源:@alexandr_wang · x.com