可引导的多元主义:基于少样本比较回归的多元对齐
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
AI总结:
提出基于少样本比较回归的可引导多元主义模型,通过上下文学习和推理适应多样用户偏好,在多元对齐任务中优于基线方法。
AI中文摘要:
大型语言模型(LLMs)目前通过人类反馈强化学习(RLHF)等技术进行对齐。然而,这些方法使用标量奖励,只能平均反映用户偏好。多元对齐则寻求在一系列属性上捕捉多样的用户偏好,超越仅关注有用性和无害性。为此,我们提出了一种基于少样本比较回归的可引导多元主义模型,能够适应个体用户偏好。我们的方法利用上下文学习和推理,基于一组细粒度属性,比较响应选项并做出对齐选择。为评估我们的算法,我们还通过改编Moral Integrity Corpus(MIC)和HelpSteer2数据集,提出了两个新的可引导多元主义基准,分别展示了我们的方法在价值对齐决策和奖励建模中的适用性。我们的少样本比较回归方法具有可解释性,兼容不同属性和LLMs,同时优于多个基线和最先进方法。我们的工作为多元对齐提供了新见解和研究方向,使LLMs的使用更公平、更具代表性,并推动了伦理AI的前沿发展。
英文摘要:
Large language models (LLMs) are currently aligned using techniques such as reinforcement learning from human feedback (RLHF). However, these methods use scalar rewards that can only reflect user preferences on average. Pluralistic alignment instead seeks to capture diverse user preferences across a set of attributes, moving beyond just helpfulness and harmlessness. Toward this end, we propose a steerable pluralistic model based on few-shot comparative regression that can adapt to individual user preferences. Our approach leverages in-context learning and reasoning, grounded in a set of fine-grained attributes, to compare response options and make aligned choices. To evaluate our algorithm, we also propose two new steerable pluralistic benchmarks by adapting the Moral Integrity Corpus (MIC) and the HelpSteer2 datasets, demonstrating the applicability of our approach to value-aligned decision-making and reward modeling, respectively. Our few-shot comparative regression approach is interpretable and compatible with different attributes and LLMs, while outperforming multiple baseline and state-of-the-art methods. Our work provides new insights and research directions in pluralistic alignment, enabling a more fair and representative use of LLMs and advancing the state-of-the-art in ethical AI.