AI 导读
Miso One 发布,这是一个 8B 参数的开源权重文本转语音模型,支持从短样本一键克隆声音,延迟 110ms。模型权重已在 GitHub 免费开放,可自行托管,音频数据不必离开本机,也无需 API;官方在发布中表示 API 访问即将推出,并提供了可直接试听的 demo。
正文
Miso One is live: an open-weights voice model built to sound like a real person reading, with actual warmth and pacing where most TTS still goes flat.
8B params, free on GitHub, with one-shot voice cloning from a short sample at 110ms latency.
Self-host it and your audio data never leaves your machine. No API needed, no lock-in.
Type any line into the demo and hear it before you clone the repo.
Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of latency. We’ve open-sourced the model weights, with API access coming soon. Hear how Miso One sounds in the thread below. Video在 X 查看被引用的帖子
来源:@kimmonismus · x.com