发表机构
University of Maryland; Microsoft(马里兰大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Waldo基准,用于评估跨语言知识差异下大语言模型的查询语言偏好,发现知识冲突时模型偏向查询语言文档,并探索了注意力头消融和LoRA训练两种缓解方法。
AI 中文摘要
大型语言模型越来越多地作为跨语言知识密集型信息检索任务的接口,通过综合多语言证据来提供服务。先前的研究表明,它们常常表现出查询语言偏好——即倾向于使用查询语言所撰写的来源——但这类研究大多在跨语言可获得等价知识的背景下进行考察。然而,当不同语言的来源对同一事实提供不完整或不一致的描述时,这种偏见就变得至关重要,因为用户所接收到的信息取决于模型选择使用的来源。为了刻画这种跨语言知识差异下的查询语言偏好,我们引入了Waldo,一个基于维基百科构建的多语言问答(QA)基准。Waldo包含12K个针对知识差距的问答对,即某一事实在一种语言中可得但在另一种语言中缺失,以及知识冲突,即不同语言版本对同一事实提供相互矛盾的描述。我们对五种语言下的八个模型进行了评估,发现当一种语言版本仅仅缺乏相关事实时,模型通常会使用另一种语言的证据,而不管查询语言如何。然而,在相互矛盾的描述下,模型的回答强烈倾向于查询语言中的文档,导致语义等价的查询因用户语言不同而引发不同的回答。最后,我们探索了两种可能缓解知识冲突下这种偏好的方法:一种机制性干预,即消融与查询语言偏好相关的注意力头,以及基于LoRA的训练,该方法将偏好差距减少了高达61.5%。
英文摘要
Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.
Comments43 pages, 6 figures