arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02935cs.CL

字符象似性与任意性:阿拉伯语自然语言处理视角

Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective

Dorieh Alomari, Irfan Ahmad, Maged S. Al-shaibani

首次发表
浏览论文内容

中文总结 AI 辅助

本研究从NLP视角探究阿拉伯语字符形式-功能关系,对比加点、无点及随机重映射的阿拉伯语,发现随机重映射可在降本减存的同时保持良好性能,表明阿拉伯语字符形式-功能关系具任意性。

中文摘要 AI 辅助

阿拉伯语文字使用28个字母,其中许多字母共享相同的基础形状(rasm),仅通过点的位置加以区分。由于早期阿拉伯语手稿书写时未加点却仍可被理解,因此去除点为检验这些视觉区分是否具有功能必要性提供了自然的测试场景。现有研究已表明,无点阿拉伯语仍可保持可读性,且能有效应用于自然语言处理(NLP),但目前仍不清楚这种成功是否依赖于保留原始的rasm分组,还是对相同简化rasm集合进行任意但一致的重映射也能达到可比性能。针对该问题,本研究对比了标准加点阿拉伯语、无点阿拉伯语,以及受限于相同19个无点rasm的任意字符重映射。研究在词级和字符级分词下生成了2000种随机重映射,并选取了熵值最高和最低的四种代表性映射。这些表示在语言建模、文本分类、序列标注、机器翻译以及原始文字恢复任务中接受评估。结果显示,无论是保留原始字符区分还是保留传统基于rasm的分组,对于实现良好的NLP性能都并非必要;随机重映射可在保持竞争力性能的同时,减少词汇量、未登录词(OOV)率、模型规模和训练成本。这些发现表明,从NLP视角来看,阿拉伯语字符的形式-功能关系在很大程度上是任意的:模型更多依赖稳定的分布结构,而非字母形式的视觉象似性。

英文摘要

Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscripts were written without dots yet remained interpretable, dot removal offers a natural test of whether these visual distinctions are functionally necessary. Prior work has shown that dotless Arabic can remain readable and effective for natural language processing (NLP), but it remains unclear whether this success depends on preserving the original rasm groupings or whether arbitrary but consistent remappings to the same reduced rasm set can achieve comparable performance. We address this question by comparing standard dotted and dotless Arabic with arbitrary character remappings constrained to the same 19 undotted rasms. We generated 2,000 random remappings under word- and character-level tokenization and selected four representative mappings with the highest and lowest entropy values. These representations were evaluated across language modeling, text classification, sequence labeling, machine translation, and restoration to the original script. The results show that neither preserving original character distinctions nor retaining traditional rasm-based groupings is necessary for strong NLP performance. Random remappings achieve competitive performance while reducing vocabulary size, out-of-vocabulary (OOV) rates, model size, and training cost. These findings suggest that, from an NLP perspective, Arabic character form-function relationships are largely arbitrary: models rely more on stable distributional structure than on the visual iconicity of letter forms.

↑