反事实推理用于可操控的多元价值观对齐大语言模型
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Renmin University of China(中国人民大学)
- Microsoft Research Asia(微软亚洲研究院)
- Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程研究中心,教育部)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
COUPLE通过反事实推理框架实现多元价值观对齐,解决现有方法在处理细粒度价值目标时的依赖性和优先级控制问题。
AI中文摘要:
随着大语言模型(LLMs)越来越多地融入服务于不同文化、社区和人口群体的应用中,将LLMs与超越平均原则(例如HHH)的多元人类价值观对齐变得至关重要。在心理学和社会价值理论如Schwartz的价值理论中,多元价值观由多个价值维度配以各种优先级表示。然而,现有方法在对齐此类细粒度价值目标时面临两个挑战:1)它们通常将多个价值视为独立且同等重要的,忽略了它们的相互依赖性和相对优先级(价值复杂性);2)它们难以精确控制细微的价值优先级,尤其是那些代表性不足的优先级(价值可操控性)。为解决这些挑战,我们提出了COUPLE,一种用于多元价值观对齐的反事实推理框架。它引入了结构因果模型(SCM)来特征化特征间的复杂依赖性和优先级,以及高层次价值维度与行为之间的因果关系。此外,它应用反事实推理来生成与任何期望的价值目标对齐的输出。得益于显式的因果建模,COUPLE也提供了更好的可解释性。我们在两个具有不同价值系统的数据集上评估了COUPLE,并展示了COUPLE在各种价值目标类型上优于其他基线方法。
英文摘要:
As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs with pluralistic human values beyond average principles (e.g., HHH). In psychological and social value theories such as Schwartz's Value Theory, pluralistic values are represented by multiple value dimensions paired with various priorities. However, existing methods encounter two challenges when aligning with such fine-grained value objectives: 1) they often treat multiple values as independent and equally important, ignoring their interdependence and relative priorities (value complexity); 2) they struggle to precisely control nuanced value priorities, especially those underrepresented ones (value steerability). To handle these challenges, we propose COUPLE, a COUnterfactual reasoning framework for PLuralistic valuE alignment. It introduces a structural causal model (SCM) to feature complex interdependency and prioritization among features, as well as the causal relationship between high-level value dimensions and behaviors. Moreover, it applies counterfactual reasoning to generate outputs aligned with any desired value objectives. Benefitting from explicit causal modeling, COUPLE also provides better interpretability. We evaluate COUPLE on two datasets with different value systems and demonstrate that COUPLE advances other baselines across diverse types of value objectives.