跳到正文
@testingcatalog· @testingcatalog · X·· 2026-09-03AI 评分48
AI 导读

用户可以用 [excited] 这样的方括号来引导语音方向,模型会将之前的轮次视为音频并据此调整表达方式。 TTS-2 Flash 覆盖 200+ 种语言,并包含相同的引导和克隆功能,首个音频字节延迟为 20ms,由 Inworld 在服务端测得。 输出会经过降噪处理,之前的轮次会作为上下文传入。 来试试看 👀 https://t.co/FbT3dmezlZ

正文

Users can steer voice direction using brackets such as [excited], where the model treats prior turns as audio and shifts delivery accordingly.

TTS-2 Flash covers 200+ languages and includes the same steering and cloning features, with 20ms to the first audio byte, measured server-side by Inworld.

The output gets denoised, and prior turns are passed in as context.

Test it out 👀
https://t.co/FbT3dmezlZ

来源:@testingcatalog · x.com