PTC-Bias:语音大模型音素级时间竞争偏置检索与解码后纠正
原标题:PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs
AI 摘要
西安交通大学利物浦大学的研究者提出 PTC-Bias,一个基于音素级时间竞争的两阶段上下文偏置框架,用于语音大模型(SpeechLLM)的稀有词识别。预填充阶段通过 PTC Retrieval 做帧同步音素解码与候选发音时间竞争,生成紧凑的偏置词候选表及对应语音区间;解码后 PTC Correction 在同一区间内对候选词与不匹配转写片段做二次局部竞争,实现选择性纠正。在 LibriSpeech 上,使用 Prompt-SLAM-ASR-7B 和 2000 个偏置词时,B-WER 相对 CTC-Filter 在 test-clean/test-other 上分别降低 23.4%/23.9%,U-WER 基本不变。
正文节选
PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs Abstract Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciation