arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低资源语言表示的几何特性

The Geometry of Low-Resource Language Representations

Francois Meyer, Jan Buys

arXiv 2608.23358首次发表:更新:

发表机构

University of Cape Town(开普敦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究从表示几何角度探究LLM中低资源与高资源语言的性能差异,发现低资源语言存在表示退化,提出几何正则化方法,实验表明该方法可降低退化并提升性能,为低资源语言CPT提供可行策略。

AI 中文摘要

大型语言模型(LLMs)中低资源语言与高资源语言之间的性能差距已广为人知,但仍不清楚是模型内部哪些因素导致了这些差异。本文从表示几何的角度来刻画这一差距,对30种语言的隐藏表示几何属性进行比较后发现,LLM的几何特性与语言数据可用性存在系统性关联。最一致的效应出现在最后几层,低资源语言会出现表示退化现象。为解决这一问题,本文研究了正则化项在持续预训练(CPT)中惩罚退化的效果。将9种基础LLMs单语适配到10种非洲语言的实验表明,几何正则化能有效降低CPT过程中的表示退化。对于更大规模的模型,基于余弦相似度的正则化相比普通CPT可小幅提升性能,在最具挑战性的任务上增益更稳定。研究证实,LLMs中低资源语言与高资源语言的表示几何存在可测量的差异,针对性的几何干预是改进低资源语言CPT的可行策略。

英文摘要

The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing the geometric properties of hidden representations across 30 languages reveals that LLM geometry is systematically related to language data availability. The most consistent effect is in final layers, where low-resource languages exhibit representational degeneration. To counter this, we investigate the effectiveness of regularisation terms to penalise degeneration during continued pretraining (CPT). Experiments monolingually adapting 9 base LLMs to 10 African languages show that geometric regularisation successfully reduces representational degeneration during CPT. For larger models, cosine similarity-based regularisation marginally improves performance over vanilla CPT, with more consistent gains on the most challenging tasks. We establish that the representational geometry of low- and high-resource languages in LLMs is measurably distinct, and that targeted geometric intervention is a viable strategy for improving CPT for low-resource languages.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑