arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言大语言模型中的分布感知语言神经元识别

Distribution-aware Language Neuron Identification in Multilingual Large Language Models

Minjun Kim, Inho Won, Junghun Yuk, Dongyeon Kim, Jihyo Kim, KyungTae Lim

arXiv 2609.10993首次发表:更新:

发表机构

KAIST; InnoCORE PRISM-AI Center(韩国科学技术院; InnoCORE PRISM-AI中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多语言大模型语言特异性神经元识别不足的问题,提出分布感知选择方法,利用激活分布成对重叠聚类,提升目标语言损害达4.9倍且不损其他语言。

AI 中文摘要

多语言大语言模型(mLLMs)包含一小部分对特定语言敏感的前馈神经元,通常称为语言特异性神经元。现有工作通过计算每个神经元在语言维度上被激活概率的熵来衡量语言特异性,其中当神经元的激活值为正时,该神经元被视为激活。然而,这种方法可能无法完全捕捉mLLMs的多语言特性,因为语言表示是分布式的且相互关联。我们提出分布感知语言神经元选择方法,该方法利用每种语言在完整激活范围(包括负值)上的激活分布之间的成对关系。具体而言,我们通过使用激活分布之间的成对重叠系数对语言进行聚类,来量化每个神经元的语言特异性。在两个mLLMs和两个保留语料库上,我们的识别器能更有效地隔离语言特异性因果效应,每个神经元对目标语言的损害最高提升4.9倍,同时保持非目标语言的性能不受影响。

英文摘要

Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. Existing work measures language specificity using the entropy of each neuron's language-wise probabilities of being active, where a neuron is considered active when its activation value is positive. However, this approach may not fully capture the multilingual nature of mLLMs, where language representations are distributional and mutually related. We propose Distribution-aware Language Neuron selection, which leverages pairwise relationships between per-language activation distributions over the full activation range, including negative values. Specifically, we quantify each neuron's language specificity by clustering languages using pairwise overlap coefficients between their activation distributions. Across two mLLMs and two held-out corpora, our identifier more effectively isolates language-specific causal effects, yielding up to 4.9$\times$ higher on-target language damage per neuron while preserving off-target language performance.

CommentsAccepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑