AI 导读
Muse Voice Transcribe 是来自 Muse Spark 系列的自回归多模态 LLM。 音频以 80ms 分块(12.5 Hz)处理,每块一个 token,在每个分块上模型决定是继续聆听还是输出文本。 使用词错误率和延迟奖励相结合的 RL 赋予其自适应延迟:它在难词上等待更久,在易词上更早输出,逐词权衡准确率与延迟。
正文
Muse Voice Transcribe is an autoregressive multimodal LLM from the Muse Spark family.
Audio is processed in 80ms chunks (12.5 Hz), one token each, and at every chunk the model decides whether to keep listening or emit text.
RL with combined word error rate and delay rewards gives it adaptive delay: it waits longer on hard words and commits sooner on easy ones, trading accuracy against latency word-by-word.
来源:@AIatMeta · x.com