面向汉藏语系语言的跨语言罗马字生态系统:普通话-粤语配对案例研究
Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study
浏览论文内容
中文总结 AI 辅助
本文提出汉藏语系罗马字生态系统框架,开发多种汉藏语系语言罗马字方案及配套开源基础设施,经实验验证其可提升低资源汉藏语系语音技术的迁移效果。
中文摘要 AI 辅助
本文提出汉藏语系罗马字生态系统(Sinitic Romanization Ecosystem),这是一个带有配套数字基础设施和社区驱动开源工作流的跨语言汉藏语系罗马字设计框架。该设计框架通过四项设计原则解决汉藏语系语言间缺乏系统性跨语言罗马字对齐的问题:一是语音对应原则,用相似罗马字符号表示相似发音;二是历史音系对应原则,对齐同源罗马字字符串;三是一音素一符号原则;四是基本拉丁字母使用原则,同时平衡考虑这些原则间的权衡关系。在主要配对案例研究中,我们依据该设计框架分别开发了粤语罗马字方案CantRomZJ1和普通话罗马字方案MandRomZJ1,还依据相同框架开发了包括梅县客家话、上海吴语、南京江淮官话在内的其他几种汉藏语系语言的罗马字方案。为将这些罗马字方案投入实际应用,我们开发了用于结构化罗马字存储、转换、解析、词典构建及输入法生成的开源基础设施。最后,我们基于Meta的大规模多语言语音(Massively Multilingual Speech, MMS)微调开展语音转罗马字实验以评估该设计框架。与拼音+粤拼基线相比,我们的MandRomZJ1+CantRomZJ1设置使粤语的词错误率(WER)降低7.80%,字符错误率(CER)降低10.61%。这些结果表明,跨语言罗马字对齐可提升低资源汉藏语系语音技术的迁移效果。
英文摘要
This paper proposes the Sinitic Romanization Ecosystem, a cross-lingual Sinitic romanization design framework with supporting digital infrastructure and a community-driven open-source workflow. The design framework addresses the lack of systematic cross-lingual romanization alignment among Sinitic languages through four design principles: phonetic correspondence for representing similar sounds with similar romanized symbols, historical-phonological correspondence for aligning cognate romanization strings, one-phoneme-one-symbol, and basic Latin-letter use, with a balancing consideration recognizing trade-offs among these principles. For the main paired case study, we develop CantRomZJ1 and MandRomZJ1, Cantonese and Mandarin romanization schemes following the design framework, respectively. We also develop schemes for several other Sinitic languages, including Meixian Hakka, Shanghai Wu, and Nanjing Jianghuai Mandarin, following the same design framework. To bring the romanization schemes into practical use, we develop open-source infrastructure for structured romanization storage, conversion, parsing, dictionary construction, and input-method generation. Finally, we evaluate the design framework through speech-to-romanization experiments based on Meta's Massively Multilingual Speech (MMS) fine-tuning. Compared with the Pinyin+Jyutping baseline, our MandRomZJ1+CantRomZJ1 condition reduces Cantonese WER and CER by 7.80% and 10.61%, respectively. These results suggest that cross-lingual romanization alignment can improve transfer in low-resource Sinitic speech technology.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。