Cartesia 推出 Sonic-3.6 文本转语音模型(90ms 延迟)和 Ink-2 语音转文本模型(100ms 转写延迟),称其流式速度超过其他同类产品,并在说话与聆听两类模型上均占据第一。语音栈正成为 AI 智能体执行循环的一部分,而语音智能体对延迟极为敏感,100ms 的偏差就会被明显感知,Cartesia 正同时攻向这一循环的两端。
There is a subtle architecture shift happening in voice AI.
The voice stack is becoming part of the agent's execution loop.
Cartesia is combining the listening and speaking paths around that loop.
Sonic-3.6 turns text into speech (90ms latency) and Ink-2 turns speech into text (100ms transcript latency), faster than anything else streaming.
now holds the #1 spot for both speaking and listening models.
Becasue, voice agents are one of those systems where 100ms in the wrong place is very noticeable.
Cartesia is attacking both sides of that loop at once.
来源:@rohanpaul_ai · x.com