arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TamilEOT:泰米尔语电话语音语义话轮结束检测的数据集与模型

TamilEOT: A Dataset and Model for Semantic End-of-Turn Detection in Tamil Telephone Speech

Santhoshkumar V

arXiv 2609.05631首次发表:更新:

AI 中文总结

针对泰米尔语电话语音,构建了含18,485个话轮边界的数据集TamilEOT,并微调两个音频检测器,准确率从70.30%提升至86.13%,运行时间低于150毫秒。

AI 中文摘要

语音智能体必须在每次停顿时判断用户是否已经说完。如果没有语言模型,这一决策就会退化为固定的静音超时:设置过短,智能体会打断用户;设置过长,每个话轮都要承受完整的等待时间。目前存在开源的语义话轮结束检测器,但据我们所知,没有一个覆盖南印度语言。我们发布了TamilEOT:从116段真实泰米尔语电话对话中切出的18,485个带标注的话轮边界,以及两个从Smart Turn v3微调而来的纯音频检测器。在来自30个未见通话的4,168个片段的留出测试集上,准确率从零样本的70.30%提升至83.71%(8.7 MB模型)和86.13%(21 MB模型);ROC-AUC从0.751提升至0.921。两个模型在笔记本电脑CPU上单线程运行时间均低于150毫秒。我们还报告了构建该数据集的成本。基于规则生成的标签,经盲听人工核对,在正类上正确率为95.9%,但在负类上仅为44.4%,低于随机水平,因为规则回答的问题与模型被问的问题不同。用音频大语言模型标注器替换后,测得人工一致率为97.5%,成本为5.69美元。在我们测量的所有训练手段中,只有编码器容量对结果有影响;在相同配置和随机种子下,三次运行的准确率波动范围为0.87个百分点,这是其他所有差异均无意义的基准下限。将相同的标注边界通过生产环境语音活动检测器和流式适配器重放,额外损失2.60个百分点,且7.8%的边界从未呈现给模型。数据、权重、代码及所有负面结果均已公开。

英文摘要

A voice agent has to decide, at every pause, whether the user has finished speaking. Without a model of the language that decision falls back to a fixed silence timeout: set it short and the agent interrupts, set it long and every turn pays the full wait. Open semantic end-of-turn detectors exist, but to our knowledge none covers a South Indian language. We release TamilEOT: 18,485 labelled turn boundaries cut from 116 real Tamil telephone conversations, and two audio-only detectors fine-tuned from Smart Turn v3. On a held-out split of 4,168 clips from 30 unseen calls, accuracy rises from 70.30% zero-shot to 83.71% (8.7 MB) and 86.13% (21 MB); ROC-AUC rises from 0.751 to 0.921. Both models run in under 150 ms single-threaded on a laptop CPU. We also report what building it cost. Rule-derived labels, checked against a blind human listening pass, were right 95.9% of the time on the positive class and 44.4% on the negative class, which is below chance, because the rule answered a different question than the model is asked. Replacing them with an audio-LLM labeller measured at 97.5% human agreement cost US$5.69. Of every training lever we measured, only encoder capacity moved the result; three runs at identical config and seed span 0.87 accuracy points, which is the floor below which none of our other deltas mean anything. Replaying the same labelled boundaries through the production VAD and streaming adapter costs a further 2.60 points, and 7.8% of boundaries are never surfaced to the model at all. Data, weights, code and every negative result are public.

Comments13 pages, 7 tables. Code, data and models: https://github.com/santhosh-005/tamil-eot

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑