AI 导读
用户可以用 [excited] 这样的方括号来引导语音方向,模型会将之前的轮次视为音频并据此调整表达方式。 TTS-2 Flash 覆盖 200+ 种语言,并包含相同的引导和克隆功能,首个音频字节延迟为 20ms,由 Inworld 在服务端测得。 输出会经过降噪处理,之前的轮次会作为上下文传入。 来试试看 👀 https://t.co/FbT3dmezlZ
正文
Users can steer voice direction using brackets such as [excited], where the model treats prior turns as audio and shifts delivery accordingly.
TTS-2 Flash covers 200+ languages and includes the same steering and cloning features, with 20ms to the first audio byte, measured server-side by Inworld.
The output gets denoised, and prior turns are passed in as context.
Test it out 👀
https://t.co/FbT3dmezlZ
来源:@testingcatalog · x.com