arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27753cs.CV

低资源语言能从彼此间学到什么?

What Can Low Resource Languages Learn From Each Other?

发表机构印度理工学院德里分校
查看机构详情
  • IIT Delhi(印度理工学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

Achyuth P, Kahaan Shah, Chetan Arora

首次发表
浏览论文内容

中文总结 AI 辅助

针对低资源语言OCR数据稀缺问题,提出PSMC框架,通过融合语言特定专家实现知识迁移,在10种印度文字上使词识别率平均提升约2%。

中文摘要 AI 辅助

尽管视觉语言模型(VLMs)发展迅速,但其语言覆盖范围仍主要局限于高资源语言,致使全球7000多种现存语言中的大多数被置于日益扩大的数字鸿沟的不利一侧。这种差距在光学字符识别(OCR)领域尤为显著,低资源文字缺乏传统缩放定律所需的海量数据集。我们在极端数据稀缺 regime(<1万张真实图像、<25万张合成图像)下研究OCR适配,发现常规微调策略常触及性能天花板。我们的关键发现揭示了语言特定适配中的结构低效性:专用模型的高层会分化以捕捉独特的文字细微差别,而低层则学习到冗余、高度相似的特征。基于此观察,我们提出PSMC(预训练、专用化、融合与协同训练)这一数据高效框架,利用跨文字的“迁移效应”。该方法首先从高资源基础模型中衍生出语言特定专家,再采用任务算术将这些专家融合为统一、高性能的多语言主干。对10种印度文字(支持20多种语言)的广泛评估显示,PSMC在不增加参数数量的情况下,相比单个专用模型实现了约2%的词识别率(WRR)平均提升。我们的结果表明,在融合潜在空间中进行联合训练促进了建设性的知识迁移,使所有组成文字受益,为包容性VLM开发提供了可扩展路径。源代码和数据集将在出版后发布。

英文摘要

Despite the rapid advancement of Vision-Language Models (VLMs), their linguistic reach remains largely confined to high-resource languages, leaving the majority of the world's 7,000+ living languages on the wrong side of a growing digital divide. This disparity is especially pronounced in Optical Character Recognition (OCR), where low-resource scripts lack the massive datasets required for traditional scaling laws. We investigate OCR adaptation in extreme data-scarce regimes (<10K real and <250K synthetic images), demonstrating that conventional fine-tuning strategies often reach a performance ceiling. Our key finding reveals a structural inefficiency in language-specific adaptation: while higher layers of specialized models diverge to capture unique script nuances, the lower layers learn redundant, highly similar features. Motivated by this observation, we propose PSMC (Pre-train, Specialize, Merge, and Co-train), a data-efficient framework that capitalizes on a cross-script "transfer effect". Our approach first derives language-specific experts from a high-resource base model, then employs task arithmetic to fuse these experts into a unified, high-performance multilingual back- bone. Extensive evaluation across 10 Indian scripts (supporting 20+ languages) shows that PSMC achieves a ~2% average improvement in Word Recognition Rate (WRR) over individual specialist models without increasing parameter count. Our results indicate that joint training in the merged latent space facilitates a constructive knowledge transfer that benefits all constituent scripts, providing a scalable pathway for inclusive VLM development. Source code and datasets will be released post publication.

补充信息

↑