arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对标准词嵌入属性的实证研究

An empirical investigation into the properties of standard word embeddings

Salomon Kabongo

arXiv 2607.23675首次发表:更新:

发表机构

African Institute for Mathematical Sciences (AIMS); North-West University(非洲数学科学研究所; 西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究词嵌入属性,回顾计算词嵌入的机制,研究公共领域的工具包和矩阵,通过选定实现进行实验以了解其特征

AI 中文摘要

词序列到连续向量空间的嵌入是近年来自然语言处理中最重要的进展之一。此类嵌入已应用于自动语音识别、机器翻译、情感分析等领域。本文回顾了计算词嵌入的各种机制,研究了公共领域中流行的工具包和嵌入矩阵,并通过一个或多个选定的实现进行实验,以更好地了解它们的特征。词的连续向量表示是近年来自然语言处理领域最重要的进展之一。这些表示已应用于语音识别、自动翻译、情感分析等领域。这项工作回顾了计算这些词向量的不同机制,研究了流行的在线工具包和公开可用的矩阵,并对一个或多个选定的实现进行实验,以更好地了解它们的特征。

英文摘要

The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processing in the recent past. Such embeddings have found application in areas such as Automatic Speech Recognition, Machine Translation, Sentiment Analysis and many more. This essay reviews the various mechanisms that have been proposed for the calculation of word embeddings, investigates popular toolkits and embedding matrices that are available in the public domain, and experiments with one or more selected implementations to better understand their characteristics. La représentation vectorielle continue de mots a été l'un des développements les plus importants dans le domaine du traitement automatique du langage naturel au cours des dernières années. Ces représentations ont trouvé application dans des domaines tels que la reconnaissance vocale, la traduction automatique, l'analyse des sentiments, etc. Ce travail passe en revue les différents mécanismes proposés pour le calcul de ces vecteurs de mots, étudie les kits d'outils populaires et les matrices disponibles publiquement en ligne, et expérimente avec une ou plusieurs implémentations sélectionnées pour mieux comprendre leurs caractéristiques.

CommentsAfrican Institute for Mathematical Sciences (AIMS) - South Africa, University of the Western Cape

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑