返回全部动态

超越WER:带口音对话ASR中的实体与不流畅词召回

原标题:Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

arXiv cs.CL一手来源研究质量 78

AI 摘要

该研究提出一个三阶段流水线,用于提升带口音英语对话语音识别中的命名实体和填充词召回率。方法包括:用启发式SQL过滤筛选实体密集训练数据、在Qwen2.5-Omni-3B上为印度、印尼和拉美地区分别训练LoRA适配器、并建立六类错误分类体系。结果显示实体召回率从53-55%提升至80-85%,填充词召回率从5%提升至76-86%,WER降至6-10%,在实体召回上超过Whisper和某商业ASR,并以更少参数匹配零样本30B模型。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Husain Pandey Singh Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR Abstract ASR systems optimised for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both critical for language-learning feedback. We present a three-stage pipeline for speakers from India, Indonesia, and Latin America: (1) heuristic SQL filters curating entity-rich training data at the entity density of random sampling, (2) regional LoRA adapters fine-t


发布时间:2026-09-21 12:00
抓取时间:2026-09-21 12:03
来源机构:arXiv
阅读原文arxiv.org