阿拉伯语TTS排行榜方言偏差调查:并非刷榜,而是方言差异
原标题:On the integrity of the Arabic TTS Arena leaderboard
AI 摘要
Hugging Face 团队在收到社区关于阿拉伯语 TTS Arena 排行榜与用户体验不符的反馈后展开调查。他们最初怀疑存在刷榜行为,但通过逐条听取投票音频和检查提示词,发现异常投票其实是用户使用海湾方言进行真实评测的结果。进一步分析表明,不同阿拉伯语方言下最佳模型各不相同,单一排行榜掩盖了这种差异。最终,团队保留了所有投票,并为排行榜增加了方言筛选功能,同时更新了示例句子以覆盖更多方言。
正文节选
Our community shared with us concerns that the leaderboard felt divergent from their own experience. They found models they liked using scoring low and models that weren't that good got the top ranks. They were afraid that there was some sort of benchmaxxing happening like a lot of metrics nowadays. We started investigating right away, and what we found was a lot more interesting than a simple leaderboard hack. We started by pulling every vote cast since the arena launched in March, and went thr