arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20346cs.CLcs.SDeess.AS

构建并评估面向电信客服场景的合成孟加拉语语音资源

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

Kawshik Kumar Paul, Md. Nafiul Alam Fuji

首次发表
浏览论文内容

中文总结 AI 辅助

本文构建并公开了面向电信客服场景的合成孟加拉语语音数据集,经评估其文本与音频一致性强,同时探讨了合成语音及STT评估的局限性。

中文摘要 AI 辅助

面向客户的应用所使用的语音系统通常需要覆盖特定领域的语言内容。本文提出了一种面向电信客服场景的合成孟加拉语语音数据集,该数据集包含10000组音频-文本对,时长约26.82小时,采用24 kHz采样率;预定义了训练、验证、测试划分,分别包含9000、500、500个样本。该数据集已在Hugging Face上以CC-BY-4.0许可公开发布。语音采用OmniVoice的语音克隆模式生成,使用真实女性参考录音及对应的转录文本,采用bfloat16精度、16步扩散采样,语速控制值设为1.0。除原始孟加拉语文本外,数据集还提供了专为ASR/STT训练与评估设计的标准化转录字段。本文使用经领域适配的Whisper ASR模型(由bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium微调而来)对全部10000个样本进行自动可懂度检查,并对选定样本开展人工听辨检查。评估得到平均词错误率(WER)为2.54%,平均字符错误率(CER)为0.59%,WER与CER的中位数均为0.00%。这些结果表明,在所选自动评估流程下,文本与音频具备强一致性,同时本文也讨论了合成语音及基于STT的评估的局限性。

英文摘要

Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under the CC-BY-4.0 license. The speech was generated with OmniVoice in voice-cloning mode using a real female reference recording and transcript, with bfloat16 precision, 16 diffusion sampling steps, and a speaking-rate control value of 1.0. Along with the original Bengali text, the dataset provides a normalized transcript field designed for ASR/STT training and evaluation. We report an automatic intelligibility check over all 10,000 samples using a domain-adapted Whisper ASR model fine-tuned from bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium, along with a manual listening check on selected samples. The evaluation gives an average WER of 2.54%, an average CER of 0.59%, and median WER and CER values of 0.00%. These results suggest strong text-audio consistency under the selected automatic evaluation pipeline, while the paper also discusses the limitations of synthetic speech and STT-based evaluation.

发表机构

  • Bangladesh University of Engineering and Technology (BUET)(孟加拉工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑