跳到正文
@kimmonismus· @kimmonismus · X·· 2026-06-04AI 评分52
AI 导读

Miso One 发布,这是一个 8B 参数的开源权重文本转语音模型,支持从短样本一键克隆声音,延迟 110ms。模型权重已在 GitHub 免费开放,可自行托管,音频数据不必离开本机,也无需 API;官方在发布中表示 API 访问即将推出,并提供了可直接试听的 demo。

正文

Miso One is live: an open-weights voice model built to sound like a real person reading, with actual warmth and pacing where most TTS still goes flat.

8B params, free on GitHub, with one-shot voice cloning from a short sample at 110ms latency.

Self-host it and your audio data never leaves your machine. No API needed, no lock-in.

Type any line into the demo and hear it before you clone the repo.

引用Aoden Teo (@AodenTeoMT)@AodenTeoMT
Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of latency. We’ve open-sourced the model weights, with API access coming soon. Hear how Miso One sounds in the thread below. Video
在 X 查看被引用的帖子

来源:@kimmonismus · x.com