arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

复合词释义基于类比

Compound interpretation is based on analogy

Tian Shen, Harald Baayen

arXiv 2610.01688首次发表:更新:

发表机构

Northwest University; University of Tübingen(西北大学; 蒂宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出复合词类比模型(CAM),通过成分嵌入与家族平均位移向量预测复合词嵌入,在普通话复合词上优于CAOSS模型,并验证了其认知合理性,表明复合词理解基于局部类比泛化。

AI 中文摘要

如何从成分意义最佳地预测复合词意义,仍是计算词汇语义学模型中的核心问题。比较不同的计算模型,为评估在复合词理解过程中语义信息组合方式的不同理论解释提供了途径。我们提出一个新模型——复合词类比模型(Compound Analogy Model, CAM),该模型通过将复合词的成分嵌入向量相加,并加上两个成分的复合词家族的平均位移向量,来预测复合词的嵌入向量。该模型无需参数,利用了语义空间中的局部类比结构。我们在普通话复合词上,将CAM与CAOSS模型进行了评估。在训练数据和保留数据上,CAM的预测准确率均持续高于CAOSS,但三字复合词除外,因为此类复合词的类比泛化受到成分家族规模较小以及两个成分家族规模显著失衡的制约。当评估基于按频率划分的训练-测试分割(这种分割更接近从熟悉复合词到新颖复合词的泛化情形)时,CAM的优势依然保持。为评估两个模型的认知合理性,我们进一步检验了模型导出的语义度量是否能预测双字复合词的视觉词汇决策潜伏期。与CAOSS模型导出的预测因子相比,CAM导出的预测因子对反应潜伏期提供了更好的预测。这些发现表明,复合词意义更宜被描述为局部类比泛化,而非学习到的全局线性变换的应用,并证明类比语义结构为复合词理解提供了认知上合理的基础。

英文摘要

How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑