跳到正文
@TencentHunyuan· @tencenthunyuan · X·· 2026-06-08AI 评分49
AI 导读

腾讯混元联合 SJTU、SII、NTU 等机构推出 MMAE——首个面向语音与音频编辑的综合评测基准,要求 AI 理解现有音频并按自然语言指令精确修改,而非仅生成音频。该基准含 2,000 条高保真样本、17,741 条细粒度评分项、7 种模态设置、6 级任务复杂度与 8 类操作类型,当前模型 Exact Match Rate(EMR)不足 5%。

正文

Can AI truly edit audio, not just generate it? 🎧

Tencent Hy, in collaboration with SJTU, SII, NTU, TJU, ZODA, PKU, FDU, and other collaborators, introduces MMAE.

MMAE--A Massive Multitask Audio Editing Benchmark, is the first comprehensive evaluation benchmark for speech and audio "Banana🍌"

Instead of simply requiring the AI to "generate" audio, it demands that the AI understand an existing audio clip and precisely modify it according to natural language instructions—altering what needs to be changed while leaving the rest untouched.

Current models show an Exact Match Rate (EMR) below 5%, revealing a major gap in reliable audio editing.

MMAE includes:

✅ 2,000 high-fidelity samples from real-world scenarios

✅ 17,741 fine-grained rubric evaluation items

✅ 7 modality settings across sound, music, speech and their mixtures

✅ 6 task complexity from basic modifications to multi-hop reasoning and multi-round editing

✅ 8 operation types across local and global granularities

How to use:

arXiv: arxiv.org/abs/2606.07229

GitHub: github.com/ddlBoJack/MMAE

HuggingFace: huggingface.co/datasets/BoJa…

Demo: piped.video/6At5nTWhlXI

Video

来源:@TencentHunyuan · x.com