发表机构
International Institute of Information Technology(国际信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文构建了六种印度语言的词语相似度数据集,通过翻译和重新标注英语数据集,并为三种语言提供词表示模型的基线评估。
AI 中文摘要
随着词表示的出现,词语相似度任务作为评估表示质量的一种度量方式正变得越来越流行。在本文中,我们提出了六种印度语言——乌尔都语、泰卢固语、马拉地语、旁遮普语、泰米尔语和古吉拉特语——的人工标注单语词语相似度数据集。这些语言是除印地语和孟加拉语之外全球使用最广泛的印度语言。对于这些数据集的构建,我们的方法依赖于对英语词语相似度数据集的翻译和重新标注。我们还通过在新创建的词语相似度数据集上评估最先进的技术,为乌尔都语、泰卢固语和马拉地语提供了词表示模型的基线分数。
英文摘要
With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian languages - Urdu, Telugu, Marathi, Punjabi, Tamil and Gujarati. These languages are most spoken Indian languages worldwide after Hindi and Bengali. For the construction of these datasets, our approach relies on translation and re-annotation of word similarity datasets of English. We also present baseline scores for word representation models using state-of-the-art techniques for Urdu, Telugu and Marathi by evaluating them on newly created word similarity datasets.