NVIDIA 研究人员提出一种前端-后端架构,让全双工语音模型具备工具调用能力,工具调用召回率 92.0% 至 97.2%,拒绝无关调用准确率 81.2%。该架构让语音前端学习输出委托 token,将流式转录转发给文本后端 LLM 执行工具调用,并通过轻量 prefill-and-repeat 机制回传结果再由流式 TTS 播报。
NVIDIA research papers are on fire recently!
Here is another interesting paper where they give full-duplex speech models tool calls.
(bookmark it)
Commercial duplex voice models complete 31 to 51 percent of grounded customer-service tasks under clean conditions.
Text agents like GPT-5 reach 85 percent on the same tasks in text mode.
Most of what a voice agent loses, it loses in the speech pipeline.
The fix routes the decision out of the speech model.
The duplex frontend learns to emit a delegation token, forwards streaming transcripts to a text backend LLM for the tool call, and receives the result through a lightweight prefill-and-repeat mechanism before streaming TTS speaks it.
Tool-call recall runs 92.0 to 97.2 percent with 81.2 percent accuracy at rejecting irrelevant calls. Turn-taking rate, streaming ASR word error rate and spoken-language intelligence all stay at the no-tool-call baseline.
Paper: https://t.co/Ltzejtx5YC
Chat with Paper: https://t.co/V3Sx33Yrta
来源:@omarsar0 · x.com