arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12278cs.CLcs.AIcs.CY

结构性沉默:当AI基础设施使代表性不足语言的使用者处于不利地位时

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Avijit Roy, Proma Roy

首次发表
浏览论文内容

中文总结 AI 辅助

本文以孟加拉语为例,指出AI基础设施的结构性缺陷使代表性不足语言使用者处于不利地位,提出数据集稀缺是结构性障碍,应采用离线优先设计以减少相关不平等。

中文摘要 AI 辅助

面向教育和语言支持的人工智能工具,日益被视为对资源不足社区接入鸿沟的可扩展解决方案。然而,支撑这些工具的基础设施,包括训练语料库、分词方案、评估基准和部署架构,在模型训练前就可能系统性地使代表性不足语言的使用者处于不利地位。本文以全球使用最广泛的语言之一孟加拉语为例,聚焦低连通性环境下的AI辅助教育,考察这些结构性障碍。我们识别出四个相互关联的失效点:一是严重的网络存在缺口,尽管孟加拉语使用者占全球人口近4%,但其内容仅占全球网络内容的不到0.5%;二是主要多语言语料库中,英语与孟加拉语的训练词元存在67:1的缺口;三是孟加拉语的音节文字带来的分词惩罚,因词元产出率更高而加剧了数据缺口;四是连通性排斥,农村地区个人互联网普及率为36.5%,而城市地区为71.4%。这些失效反映了长期以来资源分配决策、机构优先级和设计默认值,在主流AI开发中未将代表性不足语言置于中心位置。我们认为,数据集稀缺应被理解为结构性障碍而非孤立的技术局限,且离线优先设计应被视为一种面向公平的基础设施策略。最后,本文提出了语言学与AI研究的方向,旨在减少这些结构性不平等。

英文摘要

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.

发表机构

  • John Jay College of Criminal Justice, CUNY(纽约城市大学约翰杰刑事司法学院)
  • The City College of New York, CUNY(纽约城市学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑