跳到正文
@TencentHunyuan· @TencentHunyuan · X·· 26 天前AI 评分63
AI 导读

腾讯混元发布开源基础模型 AuK,用于统一的语音生成与编辑,输入为自然语言指令加参考音频,一个接口覆盖零样本 TTS、指令控制生成、内容编辑、耳语转换、去口音、音色/风格/情感编辑、语速与音高控制、增强、去噪以及多说话人和音乐分离。同步发布 AuK-Flash,为 4 步推理,在同等条件下约快 4.5 倍。代码、权重和 demo 已上线,官方同时给出论文与 GitHub 链接。

正文

🚀 AuK is officially here. Nano banana🍌 for audio

An open-source foundation model for unified speech generation and editing.
Natural-language instructions + reference audio. One interface.
Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation.
Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions.
Code, weights, and demo are live. Try it and share your feedback.

🤗 Paper & upvote: https://t.co/zEveUsuJRF
⭐ GitHub & star: https://t.co/72m9Msk8Yv

来源:@TencentHunyuan · x.com