IndicTriMix:开发三语代码混合的语言识别数据集与模型
IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing
- Sardar Vallabhbhai National Institute of Technology Surat(萨达尔·瓦拉巴伊国家理工学院苏拉特校区)
- TCS Research Hyderabad(塔塔咨询服务公司海得拉巴研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出IndicTriMix,通过微调MuRIL和XLM-RoBERTa模型,利用平行句生成代码混合数据,构建三语语言识别基准,有效实现标记级语言识别。
AI中文摘要:
在代码混合文本中的语言识别,主要出现在社交媒体中,当用户在单个话语中频繁切换多种语言时,这一任务变得至关重要。准确识别代码混合标记的语言成为迫切需求。传统语言识别模型专为单语文本设计,不适合代码混合环境中的标记级语言识别。我们将此任务表述为序列标注问题,并微调了适合印度语言的基于上下文的Transformer模型MuRIL和XLM-RoBERTa。我们在三种不同的数据配置(印地语、古吉拉特语和孟加拉语)上评估这些系统,以预测单个标记的语言标签。我们发布了一个用于代码混合标记语言识别的基准,包含人工标注的测试集。我们提出了两种使用三种语言平行句生成代码混合数据的方法。训练后的模型证明了上下文嵌入在多语言社交媒体文本中标记级语言识别的有效性。为了可复现性和促进未来研究,我们公开发布了微调后的模型。
英文摘要:
Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as a sequence labeling problem and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages. We evaluate these systems on three different data configurations (Hindi, Gujarati, and Bengali) to predict language labels for individual tokens. We release a benchmark for language identification in code-mixed tokens with manually annotated test sets. We propose two approaches of code-mixed generation using parallel sentences of three languages. The trained models demonstrate the effectiveness of contextual embeddings for token-level language identification in multilingual social media text. For reproducibility and to facilitate future research, we publicly release our fine-tuned models.