跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 15 天前AI 评分33
AI 导读

Cartesia 创始人认为,语音模型的难点不在"语音"本身,而在于持续增长的序列中维持连续状态,这或许才是更合适的抽象。他们从序列建模而非语音研究切入该领域,因此重点押注状态空间模型。

正文

One of the stranger assumptions in AI is that the same model architecture should work equally well for thinking and interacting.

A voice model has a weird job, it cannot just produce the right answer. It has to keep up with a person while the sequence keeps getting longer.

In voice AI "speech" may not the right abstraction for the hard part, rather continuous state could be it. Here, Cartesia's founder talking how they came into the space from sequence modeling rather than speech research, which is probably why they focused so heavily on state space models.

---
(Full video on “The Neon Show” YT channel, link in comment)

来源:@rohanpaul_ai · x.com