僧伽罗语连音切分的一种字符级神经方法
A Character-Level Neural Approach to Sinhala Sandhi Splitting
浏览论文内容
中文总结 AI 辅助
针对僧伽罗语连音切分任务,本文提出基于SandhiLex的字符级序列到序列方法,采用双向LSTM编码器与单向LSTM解码器,在困难的词汇化子集上取得68.40%精确匹配准确率,并建立首个神经基准。
中文摘要 AI 辅助
僧伽罗语连音切分旨在恢复隐藏在音系合并表面形式中的组成词或语素。该任务对僧伽罗语自然语言处理至关重要,因为连音会掩盖词汇边界,但此前尚无已发表的研究为僧伽罗语连音切分建立神经基准。我们基于SandhiLex开展了一项字符级序列到序列研究,使用原生僧伽罗语Unicode输入,并评估了循环编码器-解码器模型在词缀型及更复杂的词汇化、派生和词源型连音上的表现。核心挑战在于词汇化、派生和词源型连音这一困难子集,我们最好的模型——双向LSTM编码器配合单向LSTM解码器——仅达到68.40%的精确匹配准确率(字符级准确率为82.08%),远低于在更规则的词缀型子集上取得的94.00%。消融实验表明,双向编码是性能的最大贡献者,而原生僧伽罗语文字相比罗马化输入能提升精确匹配准确率。定性分析显示,许多错误是涉及边界相邻字符或合理但不正确的音系替换的近似失误。这些结果为僧伽罗语连音切分建立了经验基线,并指出数据规模、连音类型条件化以及基于注意力的解码是未来工作的主要方向。
英文摘要
Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-match accuracy (82.08\% character-level accuracy), well below the 94.00\% achieved on the more regular affixational subset. Ablations show that bidirectional encoding is the largest contributor to performance, while native Sinhala script improves exact match accuracy over romanized input. Qualitative analysis indicates that many errors are near misses involving boundary adjacent characters or plausible but incorrect phonological substitutions. These results establish an empirical baseline for Sinhala Sandhi splitting and identify data scale, Sandhi type conditioning, and attention-based decoding as the main directions for future work.
发表机构
- Informatics Institute of Technology(信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。