arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言可解释性方法的系统比较揭示了各向异性驱动的失败

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

Oskar Holmström, Marcel Bollmann, Marco Kuhlmann

arXiv 2609.04819首次发表:更新:

发表机构

Linköping University(林雪平大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究系统比较4种共享度量指标在21个多语言基础模型中的表现,发现指标分歧源于表征各向异性,控制变量后仅ILO与跨语言迁移高度相关,推荐将ILO作为主要共享度量并结合各向异性诊断。

AI 中文摘要

多语言语言模型会形成共享的跨语言表征,各类可解释性方法声称可量化这种共享程度。这些方法大多是独立开发的,当它们出现分歧时,尚不清楚这种分歧是反映模型的属性还是测量的人为结果。我们比较了来自5个家族的21个基础模型(参数规模为1.25亿至140亿)的4种共享度量指标(CKA、ANC、每token的GMM主导度以及ILO),并将每种指标与5个下游任务的跨语言迁移表现进行关联分析。我们发现,这些指标对模型跨语言共享程度的量化结果存在差异,且这种分歧源于各向异性——即表征倾向于聚集在嵌入空间的狭窄锥体内的特性。在控制模型规模、家族以及每个任务的变化后,仅ILO与跨语言迁移的相关性(斯皮尔曼ρ=0.90)仍然显著。因此,我们建议将ILO作为主要的共享度量指标,同时报告各向异性诊断结果。

英文摘要

Multilingual language models develop shared cross-lingual representations, and various interpretability methods claim to quantify this sharing. These methods have been developed largely in isolation, and when they disagree, it is unclear whether the disagreement reflects a property of the model or an artifact of the measurement. We compare four sharing metrics (CKA, ANC, GMM dominance per token, and ILO) across 21 base models from five families (125M-14B parameters) and correlate each with cross-lingual transfer on five downstream tasks. We find that the metrics differ in their quantification of cross-lingual sharing in these models and suggest that the disagreement traces to anisotropy, the tendency of representations to cluster in a narrow cone of the embedding space. Only ILO's correlation with cross-lingual transfer (Spearman's $ρ= 0.90$) survives controls for model size, family, and per-task variation. We therefore recommend ILO as the primary sharing metric, to be reported alongside anisotropy diagnostics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑