arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30984cs.CL

THA:面向高棉语的加权有限状态文本规范化和逆文本规范化

THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer

Seanghay Yath

首次发表
浏览论文内容

中文总结 AI 辅助

针对高棉语缺乏开源文本规范化工具的问题,提出基于加权有限状态转换器的Tha工具包,实现整行分词分类与音节边界约束,在谷歌测试集上取得高准确率并开源。

中文摘要 AI 辅助

语音合成需要将书面文本转换为口语形式,而语音识别输出则需要相反的转换。对于高棉语,这两个方向都没有维护良好的开源工具,而且高棉文字使这两项任务更加困难:单词之间没有空格分隔,数字词出现在普通单词内部。我们提出了Tha,一个基于加权有限状态转换器构建的高棉语文本规范化和逆文本规范化工具包。它通过一次最短路径搜索对整行文本进行分词和分类,第二个转换器拒绝高棉语音节内部的标记边界。在谷歌的高棉语测试套件上,Tha在所有274个基数词上与参考结果一致,仅有一个拼写变体差异;在2906个真实TTS提示中,它重写的158个句子中有153个是正确的。Tha在Apache 2.0许可证下开源。

英文摘要

Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur inside ordinary words. We present Tha, a Khmer text normalization and inverse text normalization toolkit built from weighted finite-state transducers. It segments and classifies a whole line in one shortest-path search, and a second transducer rejects token boundaries inside a Khmer syllable. On Google's Khmer test suite, Tha agrees with the reference on all 274 cardinals up to one spelling variant, and on 2,906 real TTS prompts, 153 of the 158 sentences it rewrites are correct. Tha is open source under the Apache 2.0 license.

发表机构

  • Digital Government Committee(数字政府委员会)

机构由 AI 辅助整理,请以论文原文为准。

↑