Cartesia 推出 Sonic-3.6 文本转语音模型(90ms 延迟)和 Ink-2 语音转文本模型(100ms 转录延迟),两者均以流式速度领先,并同时占据语音生成与识别模型榜首。其思路是将语音栈纳入智能体的执行循环,把听说两条路径围绕该循环整合,因为语音智能体对 100ms 级延迟极为敏感。
There is a subtle architecture shift happening in voice AI.
The voice stack is becoming part of the agent's execution loop.
@cartesia is combining the listening and speaking paths around that loop.
Sonic-3.6 turns text into speech (90ms latency) and Ink-2 turns speech into text (100ms transcript latency), faster than anything else streaming.
now holds the #1 spot for both speaking and listening models.
Becasue, voice agents are one of those systems where 100ms in the wrong place is very noticeable.
Cartesia is attacking both sides of that loop at once.
super interesting work by @krandiash and team.
来源:@rohanpaul_ai · x.com