arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态视觉-语言编码器在低资源语言上的失效位置

Where Do Multilingual Vision-Language Encoders Fail on Low-Resource Languages?

Donghoon Han, SungHyun Moon, Aidyn Zhakatayev, Junghun Cha, SeungJae Lee

arXiv 2608.30725首次发表:更新:

发表机构

Dnotitia Inc.(Dnotitia公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对多模态视觉-语言编码器在低资源语言上的检索性能差距,发现其源于序列结束符隐藏状态的语言轨迹随深度发散,通过调整序列结束符可显著提升低资源语言检索性能,且不损害高资源语言表现。

AI 中文摘要

现有多模态视觉-语言编码器可在单个模型中覆盖数百种语言,但在两个最先进的实例中,低资源语言(LRL,如斯瓦希里语)的检索性能比高资源语言(HRL,如英语)低30个百分点以上。本文探究该性能差距在训练后的编码器中所处的位置。先前关于模态差距和跨语言子空间的研究表明,输出端的线性语言方向会排挤与对齐相关的几何结构,本文对此提出质疑:LEACE将线性语言分类器的准确率从99%以上降至接近随机水平,迭代INLP则将其降至37%-50%,而低资源语言的检索性能仅在±1.5个百分点内波动,所有层级均值在2.2个百分点内,与随机对照实验结果一致,说明线性偏差是一种症状而非原因。相反,对齐的因果因素位于编码器的前向路径中:序列结束符(EOS)的隐藏状态的每语言轨迹随深度变化而发散。在投影器前三个块处,将EOS替换为其对应的英语值,可使某编码器的斯瓦希里语检索性能从22.1%提升至69.1%,另一编码器也出现相同结果;三项对照实验排除了位置聚合同义反复和英语特异性的影响。在训练时,通过构建前端主干网络将每种语言的投影拉向平行内容的质心,可验证该诊断,在低资源语言XM3600检索任务(1000张图像子集)上提升9.6/17.1个百分点,在另外三个基准任务中也取得一致提升,同时保留高资源语言的性能。

英文摘要

Recent multilingual vision--language encoders cover hundreds of languages in a single model, yet on two state-of-the-art instances retrieval on low-resource languages (LRL; e.g. Swahili) trails high-resource ones (HRL; e.g. English) by $30^+$\,pp. We ask where in the trained encoder this gap is located. Prior modality-gap and cross-lingual subspace work suggests a linear language direction at the output crowds out alignment-relevant geometry. We falsify this: LEACE drives the linear language classifier from $>99\%$ to near chance and iterated INLP to $37$--$50\%$ while LRL retrieval moves within $\pm 1.5$\,pp and all tier means within $2.2$\,pp, tracking random controls. The linear bias is a \emph{symptom}, not the cause. Instead, the alignment-causal factor lies along the encoder's forward path: the EOS (end-of-sequence) hidden state's per-language trajectory diverges with depth. Substituting the EOS with its parallel English value three blocks before the projector lifts Swahili from $22.1\%$ to $69.1\%$ on one encoder (and reproduces on the other); three controls rule out pooled-position tautology and English specificity. A front-layer trunk that pulls each language's projection toward the parallel-content centroid corroborates the diagnosis at training time, recovering $+9.6$ / $+17.1$\,pp on LRL XM3600 retrieval (1{,}000-image subset), with consistent gains across three further benchmarks while preserving HRL performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑