arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PERCEPT:用于波斯语-英语代码混合的词性标注与分析语料库

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi, Behnam Bahrak

arXiv 2608.10109首次发表:更新:

发表机构

School of Electrical and Computer Engineering, College of Engineering, University of Tehran; School of Engineering Science, College of Engineering, University of Tehran; Tehran Institute for Advanced Studies, Khatam University(德黑兰大学工程学院电气与计算机工程学院; 德黑兰大学工程学院工程科学学院; 哈塔米大学德黑兰高级研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文推出首个含通用依存关系词性标注的公开波斯语-英语代码混合语料库PERCEPT,结合LLM辅助标注框架完成6800条社交媒体帖子标注,通过分析揭示代码混合词的词性、位置分布规律,为相关NLP研究提供支撑。

AI 中文摘要

社交媒体已成为多语言交流的主要场所,用户常在单个话语中混合多种语言。尽管已为若干语言对开发了代码混合语料库,但波斯语-英语代码混合的研究仍相对不足。现有波斯语资源缺乏针对代码混合词的通用依存关系(UD)词性(POS)标注,限制了语言学分析和语法感知型自然语言处理(NLP)模型的开发。为解决这一缺口,本文推出PERCEPT——首个公开可用的大规模波斯语-英语代码混合语料库,其包含针对代码混合词的通用依存关系词性标注。该数据集包含从X、Instagram和Digikala收集的6800条帖子。本文还提出一种大语言模型(LLM)辅助的标注框架,可自动分配词性标注和文档级主题。人工评估表明,自动生成的标注与黄金标注之间具有高度一致性,证实了标注的可靠性。利用PERCEPT,本文对多个社交媒体平台上的波斯语-英语代码混合进行了首次全面语言学分析。分析显示,名词是代码混合词的主要类别,而其他词性类别的分布因平台而异。本文还发现,代码混合词的位置分布在各平台间显著一致,而触发效应在Digikala中明显更为显著。PERCEPT可在本URL处公开获取。

英文摘要

Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code-mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑