arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18194cs.CL

T-SANDHI:面向低资源中国台湾闽南语语音识别的变调感知自适应网络与解耦混合注入

T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

  • National Taiwan Normal University(国立台湾师范大学)
  • E.SUN Financial Holding Co., Ltd.(玉山金融控股股份有限公司)
  • EZAI

机构由 AI 辅助整理,请以论文原文为准。

Hung-Yang Sung, Chien-Chun Wang, Tien-Hong Lo, Yu-Sheng Tsao, Yung-Chang Hsu, Berlin Chen

AI总结:

针对中国台湾闽南语ASR中变调与底层调混淆导致的性能瓶颈,提出T-SANDHI,在冻结Whisper上解耦声学与词汇意图,通过轻量混合注入模块提升识别性能。

AI中文摘要:

在中国台湾闽南语自动语音识别(ASR)中,以往研究常将变调视为主要挑战,假设模型无法处理隐含的音韵变化。然而,我们在中国台湾闽南语上的实验表明,语音基础模型实际上能有效处理变调变化,真正的性能瓶颈源于这些变化与保留的底层调之间的局部混淆。为解决此问题,我们提出T-SANDHI,在冻结的Whisper骨干网络上显式解耦表层声学与底层词汇意图。利用由文本派生伪标签驱动的词典引导多任务学习结构,我们的轻量级混合注入模块通过动态门控整合独立的底层调与变调音韵流。在TAT-MOE语料库和两个盲测集上的广泛评估表明,这种显式解耦有效解决了声调映射混淆,在严格参数效率下优于基线。

英文摘要:

In Taiwanese Hokkien automatic speech recognition (ASR), prior studies often treat tone sandhi as a major challenge under the assumption that models fail to process implicit phonological variations. However, our experiments on Taiwanese Hokkien reveal that speech foundation models actually handle tone sandhi variations effectively, and the real performance bottleneck stems from a localized confusion between these variations and retained citation tones. To address this, we propose T-SANDHI to explicitly decouple surface acoustics from underlying lexical intent on top of a frozen Whisper backbone. Using a lexicon-guided multi-task learning structure driven by text-derived pseudo labels, our lightweight hybrid injection module integrates independent citation and sandhi phonetic streams via dynamic gating. Extensive evaluation on the TAT-MOE corpus and two blind test sets demonstrates that this explicit disentanglement effectively resolves tonal mapping confusion, outperforming baselines with strict parameter efficiency.

补充信息

↑