arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化目标语言特征:基于SAE的多语言推理引导

Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference

Hongsheng Wang, Philipp Koehn

arXiv 2608.04904首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多语言大模型性能跨语言差异问题,提出基于SAE的推理时引导方法,通过注入目标语言特征提升模型在XCOPA、XNLI、MGSM等基准上的准确率。

AI 中文摘要

多语言大语言模型在不同语言间表现出显著的性能差异,而现有适配方法通常需要参数更新和大量多语言训练数据。我们提出一种推理时的多语言引导方法,使用预训练的稀疏自编码器(SAE)识别并强化与目标语言相关的特征。利用多语言平行句,我们比较不同语言间的SAE激活,为每种目标语言选择少量与层相关的特征,将这些特征解码为引导信号并注入模型的隐藏状态,无需额外训练。在Gemma-3-12B-it上的实验显示,其在XCOPA上的平均准确率提升10.9个百分点,在XNLI上提升5.3个百分点,在MGSM上提升1.9个百分点。

英文摘要

Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require parameter updates and considerable multilingual training data. We propose an inference-time multilingual steering method that uses pretrained sparse autoencoders to identify and strengthen target-language-related features. Using multilingual parallel sentences, we compare SAE activations across languages and select a small number of layer-specific features associated with each target language. These features are decoded into steering signals and injected into the model's hidden states without additional training. Experiments with Gemma-3-12B-it show average accuracy improvements of 10.9 percentage points on XCOPA, 5.3 points on XNLI, and 1.9 points on MGSM.

CommentsCorrected an author name. No changes to the paper content

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑