SuTRA:具有词根感知的结构统一分词
SuTRA : Structurally-Unified Tokenization with Root Awareness
- Motilal Oswal Financial Services Ltd.(莫蒂拉尔·奥斯瓦尔金融服务有限公司)
- Indian Institute of Technology (IIT) Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有子词分词器忽略印度语言形态结构导致的形态破碎问题,提出SuTRA算法并发布相关数据集,在形态对齐、语义可恢复性及机器翻译性能上取得显著提升。
AI中文摘要:
现有子词分词器优化统计压缩,但忽略了形态结构,尤其是词根与词缀之间的关系。这对形态丰富的印度语言有害,这类语言的基本单位是复杂的正字法音节(aksharas)而非字母。基于频率的方法会过度拆分单词,随意分割词根和词缀,我们将这一现象称为形态破碎。我们提出SuTRA(具有词根感知的结构统一分词,Structurally-Unified Tokenization with Root Awareness),这是一种形态感知算法,可保持aksharas的不可分割性,并惩罚跨越形态边界的合并。我们还发布了一个针对印地语、马拉地语和古吉拉特语的新形态分割数据集。SuTRA减少了破碎现象,与BPE相比,在形态对齐(边界F1)上获得了+14.7%的峰值增益,在语义可恢复性(印地语)上获得了+34%的增益。这些结构增益在机器翻译中带来了+8.08 chrF2的平均改进。
英文摘要:
Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering. We propose SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries. We also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati. SuTRA reduces shattering, achieving peak gains of +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability (Hindi) over BPE. These structural gains yield an average improvement of +8.08 chrF2 in machine translation.