arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于质心干预融合的跨语言表示学习

Cross-lingual Representation Learning via Centroid Intervention Fusion

Wei Sun, Marie-Francine Moens

arXiv 2608.26357首次发表:更新:

发表机构

KU Leuven(鲁汶大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对大语言模型跨语言性能不均的问题,提出CIF框架整合多语言干预投影,在四类基准测试中提升了跨语言性能,尤其优化了低资源语言表现。

AI 中文摘要

大语言模型(LLMs)在多语言处理上表现不均,尤其在处理低资源语言时更为明显。推理时干预是一种轻量型方法,可通过在正向传播中修改LLMs生成的隐藏状态来提升跨语言迁移能力,无需更新模型参数。然而,现有跨语言干预方法通常学习从源语言到目标语言的独立投影,这限制了可扩展性,且无法实现跨语言知识共享。本文提出Centroid Intervention Fusion(CIF,质心干预融合),这是一种将多个多语言干预投影整合为单一语言共享算子的投影融合框架。在多语言常识推理、自然语言推理、事实编辑和机器翻译基准测试中,CIF在四个模型主干上的平均性能较最强的现有成对干预基准高出最多3.378个百分点,同时支持低资源语言的性能提升。代码可在指定URL获取。

英文摘要

Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low-resource languages. Inference-time intervention offers a lightweight way to improve cross-lingual transfer by modifying the hidden states produced by the LLMs during the forward pass, without updating model parameters. However, existing cross-lingual intervention methods typically learn separate projections from source to target languages, which limits scalability and prevents knowledge sharing across languages. We propose Centroid Intervention Fusion (CIF), a projection fusion framework that consolidates multiple multilingual intervention projections into a single language-shared operator. Across multilingual commonsense reasoning, natural language inference, factual editing, and machine translation benchmarks, CIF outperforms the strongest prior pairwise intervention baseline by up to +3.378 pp on average across four model backbones, while supporting performance gains for low resource languages. The code is available at https://github.com/VRCMF/CIF.git.

CommentsEMNLP 2026 (Main)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑