arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2609.21179cs.CL

并非所有不规则性都同等重要:日语形态屈折中罕见失败模式的因果隔离

Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection

Wen Zhang

AI总结:

通过正字法感知诊断和因果消融实验,发现日语过去式屈折中罕见的不规则子类(词干以/e/结尾需双写)贡献了不成比例的错误,其影响远超一般不规则性,提示评估需细粒度子类分析。

AI中文摘要:

神经形态生成系统在基准数据集上通常能达到较高的总体准确率,然而这种性能可能掩盖了集中在罕见形态子类中的系统性错误。我们提出了一种正字法感知的诊断方法,用于分析日语过去式动词屈折,将平假名不仅视为转录媒介,而且视为编码形态音韵结构的表征系统。使用两种字符级Transformer架构,在五个随机种子上进行评估,我们表明,尽管两个系统的总体准确率均超过97%,但一个结构上特定的不规则子类型——词干以/e/结尾且在过去式后缀前需要辅音双写的动词,占数据比例不足1%——却占据了不成比例的30-43%的残余错误,并对总错误贡献了约其出现频率34-48倍的份额。随后,我们从诊断转向因果隔离:受控消融实验表明,仅移除该子类型所产生的准确率提升大于移除所有不规则动词的总和。这些发现表明,神经形态学习中的错误集中并非由不规则性本身驱动,而是由极端低频形态模式与特定正字法过程之间的交互作用所致。我们认为,形态评估应纳入细粒度子类分析,并讨论了这对数据高效、发展上合理的语言模型预训练的启示。

英文摘要:

Neural morphological generation systems often achieve high aggregate accuracy on benchmark datasets, yet such performance can conceal systematic errors clustered in rare morphological subclasses. We present an orthography-aware diagnosis of Japanese past-tense verb inflection, treating hiragana not merely as a transcriptional medium but as a representational system that encodes morphophonological structure. Using two character-level Transformer architectures evaluated across five random seeds, we show that although both systems exceed 97% aggregate accuracy, a single structurally specific irregular subtype, verbs whose stems end in /e/ and require gemination before the past-tense suffix and make up fewer than 1% of the data, accounts for a disproportionate 30-43% share of residual errors and contributes roughly 34-48x its prevalence to total errors. We then move from diagnosis to causal isolation: controlled ablation experiments show that removing this subtype alone produces larger accuracy gains than removing all irregular verbs combined. These findings indicate that error concentration in neural morphological learning is not driven by irregularity per se, but by the interaction between extreme low-frequency morphological patterns and specific orthographic processes. We argue that morphological evaluation should incorporate fine-grained subclass analysis, and discuss implications for data-efficient, developmentally plausible language model pretraining.

补充信息

↑