arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高效变音符恢复的约束 CTC 解码

Constrained CTC Decoding for Efficient Diacritic Restoration

Rufael Marew, Amr Keleg, Hanan Aldarmaki

arXiv 2607.18946首次发表:更新:

发表机构

Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究阿拉伯语音转录本变音符恢复问题,提出基于 CTC 的高效非自回归方法,通过构建标注格并纳入硬约束进行解码,在相关测试集上评估,相比基线降低了变音符错误率,实现性能和效率提升。

AI 中文摘要

在这项工作中,我们致力于阿拉伯语音转录本的变音符恢复。大多数语音数据没有变音符,限制了对细粒度语音区别建模的能力。语音模态最近被探索作为补充基于文本的变音符恢复工作的一种方式。我们提出了一种基于连接主义时间分类(CTC)的高效非自回归语音到文本变音符标注方法。我们的方法在解码过程中纳入硬约束,通过从未标注的转录本构建字符级变音符标注格,并将假设限制在有效的变音符实现上。我们在古典阿拉伯语和现代标准阿拉伯语测试集(即 ArVoice 和 ClArTTS)上进行评估,与计算量更大的多模态变音符恢复基线相比,在两者中都显示出变音符错误率有统计学意义的降低,表明所提出的方法在性能和效率上都有提升。

英文摘要

In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently been explored as a way to complement text-based diacritic restoration efforts. We propose an efficient non-autoregressive approach for speech-to-text diacritization based on Connectionist Temporal Classification (CTC). Our method incorporates hard constraints during decoding by constructing a character-level diacritization lattice from an undiacritized transcript and restricting hypotheses to valid diacritized realizations. We evaluate on Classical Arabic and Modern Standard Arabic test sets (namely, ArVoice and ClArTTS) against a more computationally-complex multi-modal diacritic restoration baseline, and show statistically significant reductions in diacritic error rates in both, demonstrating that the proposed approach offers both performance and efficiency gains.

CommentsAccepted at Interspeech 2026. Includes an additional appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑