arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于Transformer的巴西葡萄牙语韵律边界分割

Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese

Rodrigo de Freitas Lima, Julio Cesar Galdino, Marcos Vinicius Treviso

arXiv 2607.07408首次发表:更新:

发表机构

ICMC; University of Lisbon(Instituto Superior Técnico; 里斯本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究巴西葡萄牙语韵律边界分割问题,基于Whisper提出SAMPA分割器,通过微调及评估不同配置,在多数据集测试中取得有竞争力的边界检测性能,模型能遵循多种线索检测韵律边界。

AI 中文摘要

自动韵律分割从声学和语言证据中识别语音单元之间的边界。尽管最近的深度学习方法在英语上取得了很好的效果,但巴西葡萄牙语(BP)的自动分割仍主要依赖基于规则或传统机器学习方法。本文提出了SAMPA,一种基于Whisper的分割器,它在转录BP语音时插入终端韵律边界的显式标记。我们在来自NURC-SP数据集的手动分割录音上对Whisper large-v3进行微调,并评估不同的训练和测试时过滤配置,包括在MuPe-Diversidades数据集上的分布外测试。SAMPA在各种设置下都取得了有竞争力的边界检测性能,最佳模型在保留测试分割上达到F1=0.731,在MuPe-Diversidades上达到F1=0.796。最后,通过n-gram和视听分析,我们表明我们的模型在检测韵律边界时遵循形态句法、语义和韵律线索。

英文摘要

Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for Brazilian Portuguese (BP) still relies mostly on rule-based or traditional machine-learning methods. This paper presents SAMPA, a Whisper-based segmenter that transcribes BP speech while inserting explicit markers for terminal prosodic boundaries. We fine-tune Whisper large-v3 on manually segmented recordings from the NURC-SP dataset and evaluate different training and test-time filtering configurations, including out-of-distribution testing on the MuPe-Diversidades dataset. SAMPA achieves competitive boundary-detection performance across settings, with the best models reaching F1=0.731 on the held-out test split and F1=0.796 on MuPe-Diversidades. Finally, through n-gram and acoustic-visual analyses, we show that our model follows morphosyntactic, semantic, and prosodic cues for detecting prosodic boundaries.

Comments6 pages, 5 figures, submitted to an IEEE conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑