arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15443cs.CL

词性的语义空间

Semantic Space of Parts of Speech

Jiří Milička, Ivan Kraus, Arnold Stanovský, Anna Vysloužilová, Barbora Štěpánková, Lenka Fárová, Vojtěch Cink, Šárka Dohnalová

首次发表
浏览论文内容

中文总结 AI 辅助

该研究利用word2vec嵌入和神经网络降维技术,分析词性分类的模糊性,通过构建三维语义空间可视化多语言词性间关系,揭示原型词与边界词。

中文摘要 AI 辅助

在欧洲语言学传统中,词性分类被视为清晰的范畴划分,这一观点也体现在语料库语言学中,即每个消歧后的词元都被分配唯一的词性(POS)。然而,所分配的类别在很大程度上取决于标注手册中提炼出的任意决策。由于一些词语在语义或典型句法中处于词性之间,且部分词性彼此间的距离比其他词性更近,词性分类本质上似乎是模糊的。我们利用word2vec嵌入分析这种模糊性,训练神经网络将其高维数据降维至与词性确定相关的三维空间。我们将数千个词映射到这个三维空间,揭示哪些是原型词、哪些处于边界,并可视化词性间的关系。本研究使用Universal Dependencies词性标注集,涉及法语、捷克语、芬兰语、俄语和英语。

英文摘要

Parts of speech categorization is understood in the European linguistic tradition as crisp categorization, which is also reflected in corpus linguistics, where each disambiguated token is assigned exactly one POS. However, the assigned categories are largely determined by arbitrary decisions distilled into annotation manuals. Since some words stand between parts of speech in their semantics or typical syntax, and some parts of speech are closer to each other than others, POS categorization seems inherently fuzzy. We analyze this fuzziness using word2vec embeddings, training a neural network to reduce their high dimensionality to three dimensions relevant for determining parts of speech. This creates a three-dimensional space onto which we map several thousand words, revealing which are prototypical and which lie on the boundaries, and visualizing relationships between parts of speech. The study uses Universal Dependencies POS tags for French, Czech, Finnish, Russian, and English.

↑