发表机构
The Harker School(哈克学校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次将历时词嵌入方法应用于梵语这一低资源古老语言,构建270万词元语料库,结合神经sandhi切分器与词形还原器,验证了21个语义变化中19个方向符合语文学证据,证明该范式可迁移至形态复杂的低资源语言。
AI 中文摘要
历时词嵌入已成为追踪语义变化的现代标准,但该方法主要在现代、高资源且分词良好的语言上得到验证。本文检验该范式是否适用于梵语——一种古老的低资源语言,其音韵融合(sandhi)、形态屈折、复合词和多义性构成了独特挑战。我构建了一个涵盖四个经典时期的270万词元语料库,使用神经字节级sandhi切分器和词形还原器恢复词边界,并在多种配置下训练各时期嵌入。为评估该系统,我从历史学术文献中整理了一个验证集,并通过锚点位移进行方向性测试。在21个可测试的语义变化中,19个移动方向与语文学证据相符(符号检验,p=0.00011)。我进一步展示了该语言所决定的配置选择,并讨论了改进机会。
英文摘要
Diachronic word embeddings have become the modern standard for tracking semantic change, yet they have been largely validated on modern, high-resource, and well-segmented languages. This paper tests whether the paradigm transfers to Sanskrit, an ancient, low-resource language whose phonological fusion (sandhi), morphological inflection, compounding, and polysemy pose a unique challenge. I assemble a 2.7M-token corpus spanning four canonical periods, recover word boundaries with a neural byte-level sandhi splitter and lemmatizer, and train per-period embeddings across configurations. To evaluate the system, I curate a validation set from historical scholarship and test recovery directionally with anchor displacement. Of 21 testable shifts, 19 move in the philologically attested direction (sign test, p=0.00011). I further show which configuration the language forces and comment on opportunities for improvement.