语音识别在街道名称上错误率高达39%,合成数据可降低60%
原标题:How speech models fail where it matters the most and what to do about it
AI 摘要
Together AI 的研究显示,语音识别系统在识别不同语言背景说话者所说的街道名称时,平均转录错误率达 39%,非英语母语者的准确率比英语母语者低 18%。他们提出了名为“跨语言风格迁移”的合成数据生成技术,用不到 1000 个训练样本将错误率相对降低 60%。该研究还引入了 SF Streets 和 US Streets 两个新基准,并估计仅旧金山出租车行业因转录错误每年造成约 210 万美元的损失。
正文节选
We demonstrate that voice recognition systems struggle to understand street name pronunciations when speakers have diverse linguistic backgrounds — with an average transcription error rate of 39% across 15 state-of-the-art models, and an 18% accuracy gap between non-English and English primary speakers. We show that a synthetic data generation technique called "cross-lingual style transfer" can reduce these errors by up to 60% relative to the base model, using fewer than 1,000 training samples.