发表机构
University of Stuttgart(斯图加特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究德语和英语名词复合词语义变化,引入组合性趋势预测任务,通过新颖数据集评估。用不同语义向量表示及时间粒度实验,发现目标复合词组合性随时间负趋势小,窄时间切片训练模型更优,且静态表示有竞争力。
AI 中文摘要
我们探讨了德语和英语名词复合词的语义变化现象,目的是研究和建模过去及随时间推移意义和组合程度的逐渐变化。为此,我们引入了组合性趋势预测任务,并根据一个新颖的数据集进行评估,该数据集包含跨越几十年历时语料库的23个德语和26个英语目标复合词的上下文组合性评级,独特地提供了每十年的评级和随时间的相应趋势。这些每十年的组合性评级使我们能够实证研究关于组合性随时间变化的未经测试的假设。除了从历时组合性注释中的实证观察,我们还使用了不同复杂度的语义向量表示进行实验,以及在历时数据上训练这些表示的几种时间粒度,每种表示类型产生约100个模型,每个模型覆盖历时语料库的不同1 - 5十年切片。与文献中假定的决定性趋势相反,我们发现目标复合词的组合性随时间仅有小的负趋势。在计算实验中,我们发现使用在历时数据的窄时间切片(单个十年或逐渐扩大的时间窗口)上训练的模型比在整个半个世纪窗口上训练的模型更符合每十年的组合性评级,后者是在语料库数据的整个一半上训练表示的普遍建模方法的类似物。此外,我们发现静态表示在组合性趋势预测任务中与上下文表示具有竞争力。
英文摘要
We explore the phenomenon of semantic change of German and English noun compounds, with the objective of investigating and modeling gradual changes of meanings and degrees of compositionality in the past and over time. To do so, we introduce the Compositionality Trend Prediction task, which is evaluated against a novel dataset of in-context compositionality ratings sampled across several decades of diachronic corpora for 23 German and 26 English target compounds, uniquely providing per-decade ratings and corresponding trends over time. These per-decade compositionality ratings allow us to investigate empirically untested hypotheses of generalized trends in compositionality over time, such as the idea that compounds should become less compositional (less transparent) over time. Beyond our empirical observations from the diachronic compositionality annotations, we perform experiments with semantic vector representations of varying complexity, as well as several temporal granularities for training these representations on diachronic data, resulting in about 100 models of each representation type, each covering a different 1--5 decade slice of a diachronic corpus. Contrary to the decisive tendency posited in the literature, we find only a small negative trend in compositionality over time in our target compounds. In our computational experiments, we find that using models trained on narrow time slices of diachronic data (single decades, or incrementally expanding temporal windows) align better with the per-decade compositionality ratings than those trained on an entire half-century window, the latter setting being an analog for the prevalent modeling approach of training representations on an entire half of a corpus' data. Additionally, we find static representations to be competitive with contextual representations in the Compositionality Trend Prediction task.
Comments40 pages, 31 tables, 10 figures. This is a pre-print under review by Computational Linguistics