发表机构
University of Groningen; LMU Munich; University of Amsterdam(格罗宁根大学; 慕尼黑大学; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多语言语言模型的跨语言一致性增强方法开展统一评估,发现后训练方法更可靠,直接分布对齐效果显著,且未对文化相关问答的响应能力造成系统性损害,为相关研究提供参考。
AI 中文摘要
多语言语言模型常对语义等价的问题在不同语言间给出不一致的答案,这促使人们开发提升跨语言一致性(CLC)的方法。然而,现有方法通常采用不同的模型、任务和协议进行评估,导致其相对优势尚不明确。本研究针对问答任务,对代表性的CLC增强方法开展统一评估,涵盖推理时干预方法与后训练方法,涉及三类模型家族和三类闭式基准。结果表明,后训练方法通常更可靠,直接分布对齐在所有模型-数据集组合中均能持续提升CLC,而其他方法对答案格式和语言覆盖广度更敏感。值得注意的是,除非源任务与目标任务的输出格式相似,否则跨域迁移效果有限。我们进一步探究CLC增强是否会损害模型在需要时做出不同响应的能力,即当被问及依赖文化的问题时。在两个文化多样的问答基准上,受控闭式评估未发现系统性性能下降,而开放式生成则出现偶尔的准确率降低,尤其针对非英语响应。本研究强调需从跨域鲁棒性和文化适配性两方面评估CLC增强,为未来后训练工作和基准开发提供参考。
英文摘要
Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear. In this work, we present a unified evaluation of representative CLC-enhancement methods for question answering, spanning inference-time interventions and post-training approaches across three model families and three closed-form benchmarks. The results show that post-training methods are generally more reliable, with direct distribution alignment consistently improving CLC across all model-dataset combinations, while other methods are more sensitive to answer format and the breadth of language coverage. Notably, cross-domain transfer is limited unless source and target tasks share similar output formats. We further investigate whether CLC enhancement hurts models' ability to respond differently *when needed*, that is, when asked culture-dependent questions. Across two benchmarks of culturally diverse question answering, we find no systematic degradation in controlled closed-form evaluation, whereas open-ended generation reveals occasional accuracy reductions, particularly for non-English responses. Our work highlights the need to evaluate CLC enhancement for both cross-domain robustness and culturally appropriate variation, informing future work in post-training and benchmark development.
CommentsPreprint. All code and datasets will be released upon publication