arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TurEngMix:用于土耳其语-英语代码混合语言识别和命名实体识别的文本语料库与基准

TurEngMix: A Text Corpus and Benchmark for Turkish-English Code-Mixed Language Identification and Named Entity Recognition

Ilayda Dogan, Phuong-Anh Nguyen-Le, Julia Mendelsohn

arXiv 2609.06963首次发表:更新:

发表机构

University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对土耳其语-英语代码混合文本,构建了TurEngMix语料库与基准,评估显示形态整合标记的错误率高,并发布资源以促进研究。

AI 中文摘要

自然语言处理系统在代码混合文本上表现不佳,尤其是在低资源语言对上。土耳其语-英语构成了进一步的挑战:它允许英语词干与土耳其语后缀结合,形成单个混合语言标记。我们引入了TurEngMix,一个包含5.5K条嘈杂的、自然发生的社交媒体帖子(486,974个标记)的语料库,这些帖子富含土耳其语-英语代码混合。从该语料库中,我们构建了一个新的土耳其语-英语基准,用于代码混合语言识别(LID)和命名实体识别(NER),包含15K个专家标注的标记。通过评估解码器LLM和微调编码器基线,我们发现单语土耳其语和英语标记被可靠地标注,但所有模型在混合语言标记上的LID和NER错误率都很高。对于形态整合的标记,GPT-4o和Qwen的NER错误率分别高出5.2倍和6.3倍。这突显了形态整合仍然是一个挑战。我们发布了语料库、标注和代码,以支持未来关于土耳其语-英语代码混合的计算和社会语言学研究。

英文摘要

Natural language processing systems underperform on code-mixed text, particularly for low-resource language pairs. Turkish-English poses a further challenge: it lets English stems combine with Turkish suffixes to form single mixed-language tokens. We introduce TurEngMix, a corpus of 5.5K noisy, naturally occurring social media posts (486,974 tokens) rich in Turkish-English code-mixing. From this corpus, we construct a new Turkish-English benchmark for code-mixed language identification (LID) and named entity recognition (NER), comprising 15K expert-annotated tokens. Evaluating both decoder LLM and fine-tuned encoder baselines, we find that monolingual Turkish and English tokens are labeled reliably, but all models have high error rates on mixed-language tokens for both LID and NER. For morphologically integrated tokens, NER error rates were 5.2x and 6.3x higher for GPT-4o and Qwen, respectively. This highlights how morphological integration remains a challenge. We release the corpus, annotations, and code to support future computational and sociolinguistic research on Turkish-English code-mixing.

CommentsAccepted to W-NUT 2026 (11th Workshop on Natural User-generated Text), co-located with EMNLP 2026. 15 pages, 4 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑