Hugging Face 与 VoiceArena 合作,在 Open ASR Leaderboard 中新增印地语和印度英语评测。该评测覆盖 4,888 名说话人、4 个说话人不重叠的数据划分,采用人们用自己手机录制的自然对话,每条音频含 16 个元数据列并设有私有留出集,还使用 OIWER 对照参考词格评分。
Speech recognition for the world’s languages needs benchmarks models cannot easily overfit.
@huggingface now has partnered with @voicearena_ai to add Hindi and Indian English to the Open ASR Leaderboard.
The evaluation design is so very interesting here:
- 4,888 speakers across 4 speaker-disjoint splits
- spontaneous conversation recorded on people's own phones
- 16 metadata columns per clip, held-out private splits
- OIWER against reference lattices where 1 written form is not enough.
A model has fewer shortcuts here: it can’t lean on familiar speakers, clean studio audio, or one rigid transcript. The score has to survive messier speech and genuinely held-out evaluation.
So the evaluation design is doing more work here than data volume.
read their full thread below, which has got so much more interesting details, on the benchmark design.
来源:@rohanpaul_ai · x.com