返回全部动态

构建并评估面向电信客服的合成孟加拉语语音资源

原标题:Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

arXiv cs.CL一手来源研究质量 78

AI 摘要

该论文介绍了一个面向电信客服场景的合成孟加拉语语音数据集,包含10,000个音频-文本对,约26.82小时,已在Hugging Face上以CC-BY-4.0许可发布。数据集使用OmniVoice的语音克隆模式生成,并提供了原始文本和归一化文本字段。通过微调的Whisper ASR模型评估,平均词错误率为2.54%,字符错误率为0.59%,表明文本-音频一致性较强。该资源旨在支持电信领域语音系统的训练和评估。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care Abstract Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under th


发布时间:2026-08-24 12:00
抓取时间:2026-08-24 12:04
来源机构:arXiv
阅读原文arxiv.org