构建并评估面向电信客服的合成孟加拉语语音资源
原标题:Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care
AI 摘要
该论文介绍了一个面向电信客服场景的合成孟加拉语语音数据集,包含10,000个音频-文本对,约26.82小时,已在Hugging Face上以CC-BY-4.0许可发布。数据集使用OmniVoice的语音克隆模式生成,并提供了原始文本和归一化文本字段。通过微调的Whisper ASR模型评估,平均词错误率为2.54%,字符错误率为0.59%,表明文本-音频一致性较强。该资源旨在支持电信领域语音系统的训练和评估。
正文节选
Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care Abstract Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under th