发表机构
University of Southern California; Information Sciences Institute(南加州大学; 信息科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出了群体对齐语言模型PALMs,通过结合心理与文化构造的理由进行偏好调优,在多维度评估中优于基准模型,且下游任务泛化能力强
AI 中文摘要
大语言模型被广泛用于模拟个体用户行为,但要忠实地代表一个群体,需要捕捉区分不同群体的价值观、信念和文化规范的系统性差异。我们推出了群体对齐语言模型(Population Aligned Language Models, PALMs),这是一组分别对齐特定群体的模型,覆盖美国、印度、巴西、法国和意大利五个国家。PALMs的构建方式是综合基于心理和文化构造的理由,并将这些理由作为群体特定对齐的偏好调优过程中的潜在监督信号。在人格、价值观与信念、文化规范和道德四个维度的评估中,PALMs始终优于包括文化专用模型在内的基准模型,在所有五个群体上的平均相对改进达到8.59%。值得注意的是,基于构造的理由的表现优于人口统计提示和基于调查的微调,这表明将偏好学习建立在心理学和文化基础上,能提供比表面层面响应分布更丰富的归纳信号。我们进一步证明,PALMs无需特定任务监督即可在下游应用中实现强泛化:在个性化奖励建模中比最佳基准模型高出5.19%,在群体模拟中高出6.34%,并在社会推理任务中表现出强迁移性。数据集和代码可在以下网址获取:this https URL
英文摘要
Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.