看见不可见之处:面向像素语言模型适配的视觉相似性
Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation
浏览论文内容
中文总结 AI 辅助
本研究以藏语为案例,针对像素语言模型适配低资源语言的问题,提出四个渲染级视觉文字相似性指标,发现正字法邻近度可增强语义迁移,且多语言预训练的PIXEL-M4与混合文字的单语言模型PIXEL的适配表现存在不对称性。
中文摘要 AI 辅助
基于像素的语言模型(LMs)通过处理文本的渲染图像替代了传统的分词器,使得跨语言迁移在很大程度上依赖于书写系统的视觉和结构属性。然而,将这些模型适配到具有复杂形态、使用独特文字书写的低资源语言的动态过程尚未得到探索。我们以藏语为案例,分析了像素语言模型的持续预训练如何受数据规模、初始文字暴露以及来自其他婆罗米系文字语言的跨语言迁移的影响。我们提出了四个渲染级指标来量化视觉文字相似性,并在三个下游任务上评估了性能。结果表明,即使在严重的数据约束下,更高的正字法邻近度也能增强语义迁移。此外,我们发现基于预训练起点存在性能不对称:多语言预训练的PIXEL-M4虽然初始性能更强,但其后续适配能力似乎受到限制,而使用混合文字的单语言模型PIXEL进行适配则在句子级任务上获得了更多收益。我们的指标和案例研究提供的实证观察,或可在类似低资源场景中使用像素模型时,为数据选择和文字适配决策提供参考。
英文摘要
Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent on the visual and structural properties of writing systems. However, the dynamics of adapting these models to low-resource languages with complex morphology and written in unique scripts are not yet explored. Using Tibetan as a case study, we analyze how continued pre-training of pixel-based LMs is influenced by data scale, initial script exposure, and cross-lingual transfer from languages written in other Brahmic scripts. We introduce four rendering-level metrics to quantify visual script similarity. We evaluate downstream performance across three tasks. Our results show that higher orthographic proximity enhances semantic transfer, even under severe data constraints. Additionally, we find a performance asymmetry based on the pre-training starting point: while multilingual pre-training PIXEL-M4 has stronger initial performance, its capacity for subsequent adaptation seems to be constrained, whereas adapting a monolingual model PIXEL with mixed scripts yields more gains on sentence-level tasks. Our metrics and case study offer empirical observations that could help inform data selection and script adaptation choices when working with pixel-based models in similar low-resource settings.
发表机构
- KU Leuven(KU Leuven(天主教鲁汶大学))
机构由 AI 辅助整理,请以论文原文为准。