Grok Voice Transcribe 2.0 发布,官方称准确率是上一版的两倍且价格不变,短语音指令错误率从 20.6% 降到 6.8%。它在 Artificial Analysis 的 32 个流式模型中准确率排第一,支持数十种语言、自动识别语言与语种切换、说话人分离标注以及最多 8 路音频独立转写,现有 API 集成无需改代码即可获得精度升级。
BREAKING: SpaceXAI has launched Grok Voice Transcribe 2.0, delivering twice the accuracy of its previous model at the same price.
• Ranks #1 for accuracy among 32 streaming models on Artificial Analysis
• Transcribes dozens of languages
• Automatically detects the language being spoken
• Handles language switching within the same recording
• Transcribes uploaded files, URLs and live audio streams
• Adds precise timestamps and confidence scores for every word
• Separates and labels different speakers at no extra cost
• Transcribes up to 8 audio channels independently
• Lets users add up to 100 key terms, including product names and medical words
• Detects when someone has finished speaking, making voice agents feel more natural
• Handles noisy calls, different accents, overlapping voices, phone numbers and email addresses
• Existing API integrations receive the accuracy upgrade without any code changes
Its error rate for short voice commands dropped from 20.6% to just 6.8%.
Pricing remains incredibly low:
• Batch transcription: $0.10 per hour
• Real-time streaming: $0.20 per hour
Atlassian found it more accurate than its previous solution and now uses it to transcribe every Loom video.
Grok Voice Transcribe 2.0 will soon become the default model in xAI’s Speech-to-Text API.
来源:@cb_doge · x.com