语料选择对依存距离估计的影响有多大?
How Much Does Corpus Choice Change Dependency-Distance Estimates?
AI总结:
该研究对比Universal Dependencies v2.18的38组同语言树库对,发现树库选择会显著影响依存距离估计的跨语言排序,仅依存长度最小化的定性规律稳定存在。
AI中文摘要:
从单一语料库得出的依存距离估计值通常被视为语言的固有属性,但这一假设尚未在独立编制的语料库间得到验证。我们采用Universal Dependencies v2.18中的38组同语言树库对,通过 concordance correlation(组内相关系数)、Bland-Altman分析及包含12种设定的多verse设计,对比了平均依存距离估计值。树库间的一致性最多仅为中等水平:将一个树库替换为另一个树库会反转近40%的成对语言排序,树库选择约占组间方差的29%。这种分歧远超过树库内部的抽样误差,且在全部12种预处理设定下均存在。不过,每个树库都证实了依存长度最小化(标准化比值低于1)。数据更支持MDD是语法、语域和标注因素的语料库条件复合体,而非稳定的语言层面参数:定性的DLM普遍性在语料库替换后依然存在,但跨语言的序数排序则不复存在。
英文摘要:
Dependency-distance estimates derived from a single corpus are routinely treated as properties of a language, yet this assumption has not been tested across independently compiled corpora. We compared mean dependency-distance estimates across 38 same-language treebank pairs from Universal Dependencies v2.18, using concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design. Cross-treebank agreement was moderate at best: substituting one treebank for another reversed nearly 40 percent of pairwise language orderings, and treebank choice accounted for roughly 29 percent of between-group variance. This disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications. Nevertheless, every treebank confirmed dependency-length minimization (normalized ratio below 1). The data are more consistent with MDD as a corpus-conditioned composite of grammatical, register, and annotation factors than as a stable language-level parameter: the qualitative DLM universal survives corpus substitution, but the ordinal cross-linguistic ranking does not.