跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 23 天前AI 评分31
AI 导读

我很喜欢 @voicearena_ai 设计的这个评估方式。 它没有简单地问一个语音模型是否比另一个"更好",而是把问题拆成任务完成度和自然度,然后用盲测成对投票分别衡量这两项。 简单的区分。但数字能告诉你的东西差别巨大。 在任务完成度上,人类和模型显然接近得多。 在自然度上,并非如此。

正文

I really like how @voicearena_ai designed this evaluation.

Instead of asking whether one voice model is simply "better" than another, they split the problem into Task Completion and Naturalness, then use blind pairwise voting to measure them separately.

Simple distinction. Huge difference in what the numbers tell you.

On Task Completion, humans and models are apparently much closer.

On Naturalness, they are not.

来源:@rohanpaul_ai · x.com