arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29371cs.CL

BanglaTurn:孟加拉语语音中话轮结束检测的基准与基于Whisper的模型

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

  • Vivasoft Limited(维瓦索夫特有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Mizbaul Haque Maruf

AI总结:

本文提出BanglaTurn语料库及基于Whisper的模型,用于孟加拉语对话语音的话轮结束检测,在测试集上准确率达84.33%,显著优于基线,并降低了假阴性率。

AI中文摘要:

本文介绍了BanglaTurn,一个用于孟加拉语对话语音中话轮结束检测的语料库,以及在该语料库上训练的模型。该语料库包含35,374个样本,每个样本为3至15秒的播客语音,通过结合说话人日志和LLM处理来标注话轮状态,每个标签随后由人工标注员进行检查。该模型将Whisper编码器与任务特定的分类头配对。在一个从保留播客中抽取的类别平衡测试集上,该模型达到了84.33%的准确率(95%置信区间80.3至88.1),而Smart-Turn v3基线的准确率为69.28%,并将假阴性率从51.57%降至7.55%,但代价是假阳性率更高。我们报告了编码器层微调、多尺度池化和INT8量化各自带来的贡献,且在CPU上的端到端延迟保持在165至191毫秒之间。

英文摘要:

This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining speaker diarization with an LLM pass, with every label then checked by a human annotator. The model pairs a Whisper encoder with task-specific classification heads. On a class-balanced test set drawn from a held-out podcast, it reaches 84.33% accuracy (95% CI 80.3 to 88.1) against 69.28% for the Smart-Turn v3 baseline, and lowers the false negative rate from 51.57% to 7.55% at the cost of a higher false positive rate. We report what encoder layer fine-tuning, multi-scale pooling and INT8 quantization each contribute, and latency stays within 165 to 191 ms end to end on CPU.

↑