arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08515cs.CLcs.AI

相同的价值观,不同的语言?从多语言探针到引导大语言模型趋向中国社会价值观

Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values

Yuemei Xu, Kexin Xu, Jian Zhou, Haoyu Lu, Yequan Wang, Aishan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建多语言对比数据集C-Voices,提出无需微调的价值观向量引导方法,发现大语言模型价值观偏好存在语言敏感性,并实现有效的跨语言价值观对齐。

中文摘要 AI 辅助

随着大语言模型(LLMs)日益融入人类社会,使其与多元社会价值观对齐已成为一个关键优先事项。然而,LLMs在不同语言中是否表现出一致的价值观偏好仍未得到充分探索,尤其是对于文化根基深厚的价值观,这些价值观比以安全为中心的原则更为抽象,难以评估和对齐。我们通过中国社会价值观(CSV)来研究这一问题,CSV是一个根植于中国文化、包含国家、社会和个人三个层面的12个维度的价值体系。我们构建了C-Voices,这是首个针对CSV的全面多语言对比探针数据集,包含六种语言、86,400个基于困境的实例,每个实例将一个符合CSV的行为与一个冲突的替代行为配对。基于C-Voices的对比探针,我们进一步提出了一种无需微调的价值观向量引导方法,该方法从隐藏状态差异中推导价值观方向,并在推理过程中选择性地干预对价值观敏感的层。在六种语言上的实验表明,CSV导向的偏好是模型依赖且语言敏感的,同一困境在不同语言中引发不同的反应。我们的方法实现了有效的CSV引导,支持价值观向量的跨语言迁移,并推广到现有的FLAMES和ValuePrism。

英文摘要

As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social Values (CSV), a value system rooted in Chinese culture and comprising $12$ dimensions across national, societal, and personal levels. We construct C-Voices, the first comprehensive multilingual contrastive probe dataset for CSV, with 86,400 dilemma-based instances in six languages, each pairing a CSV-aligned action with a value-conflicting alternative. Building on the contrastive probes of C-Voices, we then propose a fine-tuning-free value vector steering method that derives value directions from hidden-state discrepancies and selectively intervenes on value-sensitive layers during inference. Experiments on six languages show that CSV-oriented preferences are model-dependent and language-sensitive, with the same dilemma eliciting divergent responses across languages. Our method achieves effective CSV steering, supports cross-lingual transfer of value vectors, and generalizes to existing FLAMES and ValuePrism.

发表机构

  • School of Information Science and Technology, Beijing Foreign Studies University(北京外国语大学信息科学与技术学院)
  • Beijing Academy of Artificial Intelligence(北京人工智能研究院)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

↑