AI 中文总结
研究大语言模型跨语言事实不一致问题,评估零样本上下文引导、CAA、DPO四种干预策略,通过实验发现角色提示是最强干预方式,CAA对配置敏感,DPO适配器收益窄且可转移性低,表明跨语言不一致部分是选择问题,简单干预可能更优。
AI 中文摘要
尽管大语言模型(LLMs)展现出卓越的多语言流利程度,但其内部知识表示仍过度偏向高资源语言,导致跨语言事实不一致,即仅根据提示语言改变经验答案分布。我们研究能否在推理时减轻这些偏差,促使英语提示模型像以目标语言(德语、西班牙语、保加利亚语)查询时那样回答,并评估了四种干预策略:零样本上下文引导(角色提示)、通过对比激活添加(CAA)进行内部表示操纵、以及通过基于基准导出的事实数据和概念泛化数据训练的直接偏好优化(DPO)进行轻量级权重修改。为评估一致性,我们策划了一个多语言事实数据集以及一个包含文化根源查询的新颖泛化基准,以确定事实干预是否能转移到更广泛的以目标为中心的偏好上。在Gemma 3 12B Instruct上的实验表明,角色提示是总体上最强的干预方式,能平衡有效性、安全性和域外泛化。虽然CAA会导致基准不一致性大幅变化,但它对配置敏感且有知识退化风险。基于DPO的适配器提供了永久性但更窄且可转移性更低的收益。这些发现表明,跨语言不一致至少部分是一个选择问题,简单的上下文干预可能比更具侵入性的方法在稳健、可转移的对齐方面表现更好。
英文摘要
Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages. This leads to cross-lingual factual inconsistency, where they shift their empirical answer distributions based solely on the prompt language. We investigate whether these biases can be mitigated at inference time, forcing an English-prompted model to answer as if it were queried in target languages (German, Spanish, Bulgarian), and evaluate four intervention strategies: zero-shot contextual steering (persona prompting), internal representation manipulation via Contrastive Activation Addition (CAA), and lightweight weight modification via Direct Preference Optimization (DPO) trained on benchmark-derived factual data as well as conceptual generalization data. To assess alignment, we curate a multilingual factual dataset alongside a novel generalization benchmark comprising culturally rooted queries to determine whether factual interventions transfer to broader target-centric preferences. Experiments on Gemma 3 12B Instruct reveal persona prompting to be the strongest overall intervention, balancing efficacy, safety, and out-of-domain generalization. While CAA yields sharp inconsistency benchmark shifts, it is configuration-sensitive and risks knowledge degradation. DPO-based adapters offer permanent, yet narrower and less transferable gains. These findings suggest that cross-lingual inconsistency is at least partly a selection problem, and that simple contextual interventions may outperform more invasive methods for robust, transferable alignment.
Comments8 pages (21 in total), 2 figures, 4 tables. Original manuscript for a Guided Research project conducted at the Technical University of Munich, detailing the complete methodology, full data pipeline, and comprehensive experimental results. A related, condensed subset of this work was subsequently adapted and published at the StereACuLT 2026 workshop