arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

印度语言词语相似度数据集:标注与基线系统

Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems

Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, M. Shrivastava

arXiv 2609.34138首次发表:更新:

发表机构

International Institute of Information Technology(国际信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建了六种印度语言的词语相似度数据集,通过翻译和重新标注英语数据集,并为三种语言提供词表示模型的基线评估。

AI 中文摘要

随着词表示的出现,词语相似度任务作为评估表示质量的一种度量方式正变得越来越流行。在本文中,我们提出了六种印度语言——乌尔都语、泰卢固语、马拉地语、旁遮普语、泰米尔语和古吉拉特语——的人工标注单语词语相似度数据集。这些语言是除印地语和孟加拉语之外全球使用最广泛的印度语言。对于这些数据集的构建,我们的方法依赖于对英语词语相似度数据集的翻译和重新标注。我们还通过在新创建的词语相似度数据集上评估最先进的技术,为乌尔都语、泰卢固语和马拉地语提供了词表示模型的基线分数。

英文摘要

With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian languages - Urdu, Telugu, Marathi, Punjabi, Tamil and Gujarati. These languages are most spoken Indian languages worldwide after Hindi and Bengali. For the construction of these datasets, our approach relies on translation and re-annotation of word similarity datasets of English. We also present baseline scores for word representation models using state-of-the-art techniques for Urdu, Telugu and Marathi by evaluating them on newly created word similarity datasets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑