arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00325cs.CL

多语言语言模型中语言控制的潜在机制

Latent Mechanisms of Language Control in Multilingual Language Models

Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly

首次发表
浏览论文内容

中文总结 AI 辅助

本研究对比三种多语言潜在变量识别方法,引入两个多语言基准,在Gemma-2-2B和Qwen3-4B上验证FreqSel性能最优,发现语言控制潜在变量存在冗余性。

中文摘要 AI 辅助

多语言大语言模型可能会出现意外的语码转换,即在生成过程中不必要地在不同语言间切换。本文对三种识别跨层转码器中语言控制潜在变量的方法展开对比研究:基于激活值的选择(ValSel)、基于激活频率的选择(FreqSel)以及基于大语言模型生成的潜在变量标注的选择(AnnSel)。为评估这些方法识别语言控制潜在变量的有效性,本文引入两个存在语码转换的多语言基准,用于对七种语言的语言引导进行细粒度分析。通过在Gemma-2-2B和Qwen3-4B上开展针对性干预实验,研究发现三种方法均能有效操纵生成语言,其中FreqSel的整体性能最强,而AnnSel可通过明确的语言标注实现可解释的潜在变量选择。淘汰分析表明,这些方法选择的潜在变量子集互不重叠但各自具备功能,说明存在冗余性而非单一的规范语言方向。代码和数据可在该https URL获取。

英文摘要

Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in cross-layer transcoders: activation value-based selection (ValSel), activation frequency-based selection (FreqSel), and LLM-generated latent annotation-based selection (AnnSel). To evaluate the efficacy of these methods in identifying language-controlling latents, we introduce two multilingual benchmarks that exhibit code-switching for fine-grained analysis of language steering across seven languages. Through targeted intervention experiments on Gemma-2-2B and Qwen3-4B, we find that all three methods effectively manipulate generation language, with FreqSel achieving the strongest overall performance, while AnnSel offering interpretable latent selection through explicit language annotations. A knock-out analysis suggests the methods select non-overlapping but each-functional latent subsets, indicating redundancy rather than a single canonical language direction. Code and data can be found at https://github.com/rm-3284/Latent-Mechanism-Multilingual.

发表机构

  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑