arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24698cs.CL

波斯语文本自然语言处理中的标注瓶颈:作为标注稀缺语言的波斯语

The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language

MohammadHossein Mortazavi, Mostafa Salehi, Hadi Veisi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出波斯语是标注稀缺而非单纯低资源语言,通过三项定量交叉验证分析其34种文本资源,发现标注稀缺源于任务领域覆盖不均等问题,而非整体标注量不足。

中文摘要 AI 辅助

波斯语(法尔西语)在自然语言处理中常被描述为低资源语言,但这一标签将不同的资源短缺归为单一类别。本文提出,若将“标注稀缺”理解为其NLP资源生态的属性而非语言的内在属性,波斯语可更准确地被描述为标注稀缺。本文综述了截至2026年7月可用的34种代表性波斯语文本资源,并补充了三项定量交叉验证:其一,独立网络测量显示波斯语属于约20种最受关注的内容语言;W3Techs报告显示波斯语在已知内容语言的网站中占比约0.9%,而Common Crawl的CC-MAIN-2026-30数据集识别出波斯语作为主要语言的HTML页面占比为0.7039%。其二,选择性语音资源综述显示,其资源轨迹从FARSDAT延伸至近期包含数百或数千小时语音的语料库。其三,匹配的波斯语-英语比较通过Common Crawl的相对网络存在量对特定任务的标注量进行归一化,所得比率差异显著:波斯语的句法分析和新闻命名实体识别(NER)标注相对密集,而自然语言推理任务则低于网络比例基线。因此,证据不支持波斯语在标注总量上存在全球不足的简单主张,相反,标注稀缺体现为任务与领域覆盖不均、标注方案不兼容、访问与文档摩擦,以及专业领域、偏好数据和标准伊朗波斯语之外的变体的监督有限。

英文摘要

Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languages: W3Techs reports Persian on about 0.9% of websites with a known content language, while Common Crawl CC-MAIN-2026-30 identifies Persian as the primary language of 0.7039% of HTML pages. Second, a selective speech review shows a long resource trajectory from FARSDAT to recent corpora containing hundreds or thousands of hours of speech. Third, a matched Persian-English comparison normalizes task-specific annotation volumes by relative Common Crawl web presence. The resulting ratios vary sharply: Persian syntax and news NER are comparatively dense, whereas natural-language inference falls below the web-proportional baseline. The evidence therefore does not support a simple claim that Persian is globally deficient in labeled volume. Instead, annotation scarcity is expressed through uneven task and domain coverage, incompatible schemes, access and documentation friction, and limited supervision for specialist domains, preference data, and varieties beyond standard Iranian Persian.

补充信息

↑