相同的价值观,不同的语言?从多语言探针到引导大语言模型趋向中国社会价值观
Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values
浏览论文内容
中文总结 AI 辅助
本研究构建多语言对比数据集C-Voices,提出无需微调的价值观向量引导方法,发现大语言模型价值观偏好存在语言敏感性,并实现有效的跨语言价值观对齐。
中文摘要 AI 辅助
随着大语言模型(LLMs)日益融入人类社会,使其与多元社会价值观对齐已成为一个关键优先事项。然而,LLMs在不同语言中是否表现出一致的价值观偏好仍未得到充分探索,尤其是对于文化根基深厚的价值观,这些价值观比以安全为中心的原则更为抽象,难以评估和对齐。我们通过中国社会价值观(CSV)来研究这一问题,CSV是一个根植于中国文化、包含国家、社会和个人三个层面的12个维度的价值体系。我们构建了C-Voices,这是首个针对CSV的全面多语言对比探针数据集,包含六种语言、86,400个基于困境的实例,每个实例将一个符合CSV的行为与一个冲突的替代行为配对。基于C-Voices的对比探针,我们进一步提出了一种无需微调的价值观向量引导方法,该方法从隐藏状态差异中推导价值观方向,并在推理过程中选择性地干预对价值观敏感的层。在六种语言上的实验表明,CSV导向的偏好是模型依赖且语言敏感的,同一困境在不同语言中引发不同的反应。我们的方法实现了有效的CSV引导,支持价值观向量的跨语言迁移,并推广到现有的FLAMES和ValuePrism。
英文摘要
As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social Values (CSV), a value system rooted in Chinese culture and comprising $12$ dimensions across national, societal, and personal levels. We construct C-Voices, the first comprehensive multilingual contrastive probe dataset for CSV, with 86,400 dilemma-based instances in six languages, each pairing a CSV-aligned action with a value-conflicting alternative. Building on the contrastive probes of C-Voices, we then propose a fine-tuning-free value vector steering method that derives value directions from hidden-state discrepancies and selectively intervenes on value-sensitive layers during inference. Experiments on six languages show that CSV-oriented preferences are model-dependent and language-sensitive, with the same dilemma eliciting divergent responses across languages. Our method achieves effective CSV steering, supports cross-lingual transfer of value vectors, and generalizes to existing FLAMES and ValuePrism.
发表机构
- School of Information Science and Technology, Beijing Foreign Studies University(北京外国语大学信息科学与技术学院)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。