跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 20 天前AI 评分50
AI 导读

在 Final Transcript 上,Grok Voice Transcribe 2.0 在语音结束后 0.49s 实现 2.7% WER,在 AA-WER Streaming 的 33 个模型中排名第 1,较 Grok Voice Transcribe 1.0 的 3.9% WER 有所提升。它比 Muse Voice Transcribe(3.1% WER、0.16s)、带语义端点的 Cartesia Ink-2(3.4% WER、0.43s)以及 ElevenLabs Scribe v2 Realtime(3.6% WER、0.14s)更准确,但在速度上有所取舍。

正文

On Final Transcript, Grok Voice Transcribe 2.0 achieves a 2.7% WER at 0.49s after end of speech, ranking #1 of 33 models on AA-WER Streaming and improving on Grok Voice Transcribe 1.0 at 3.9% WER. It is more accurate than Muse Voice Transcribe at 3.1% WER and 0.16s, Cartesia Ink-2 with semantic endpoints at 3.4% WER and 0.43s, and ElevenLabs Scribe v2 Realtime at 3.6% WER and 0.14s, while trading some speed.

来源:@ArtificialAnlys · x.com