将大语言模型条件化于社会价值取向可提升其在序贯社会困境中的行为对齐
Conditioning LLMs on Social Value Orientation improves behavioural alignment in a sequential social dilemma
AI总结:
本研究通过将大语言模型条件化于社会价值取向,显著提升了其在蜈蚣博弈中的行为对齐,表明SVO的效用在于映射而非稳定偏好。
AI中文摘要:
大语言模型(LLMs)越来越多地被用于模拟人类决策,但其输出往往未能充分代表人类行为的异质性。我们研究了将LLMs条件化于社会价值取向(SVO,一种衡量个体相对于他人如何评价自身结果的指标)是否能更好地再现人类在序贯社会困境中的行为。以蜈蚣博弈(CG)两种变体的实验数据为参照,我们比较了八种LLMs的默认行为与根据人类样本中的SVO分布进行条件化后生成的行为。我们发现,各模型的默认策略差异显著,但总体上与参考分布相去甚远。然而,将其条件化于SVO分布会系统性地引导其策略行为,与默认情况相比,与人类参考的对齐程度最多提升了70%。在所有模型中,较高的诱导SVO值降低了停止游戏的概率,再现了人类行为中观察到的亲社会性与合作之间的关系。结合LLMs所表现出的SVO对提示和顺序效应的敏感性,这些结果表明,SVO在行为模拟中的效用并不依赖于LLMs具有稳定的社会偏好,而在于其将社会偏好映射到相应策略选择的能力。
英文摘要:
Large language Models (LLMs) are increasingly used to simulate human decision-making, yet their outputs often under-represent human behavioural heterogeneity. We investigate whether conditioning LLMs on Social Value Orientation (SVO, a measure of how individuals value their own outcomes relative to others') can better reproduce human behaviour in a sequential social dilemma. Using experimental data from two variants of the Centipede Game (CG) as reference, we compare the default behaviour of eight LLMs with behaviour generated after conditioning them on SVO profiles drawn from the human sample. We find that the default strategies vary substantially across models, but are generally distant from the reference distributions. However, conditioning them on SVO profiles systematically steers their strategic behaviour, improving alignment with the human reference by up to 70% with respect to the default. Across models, higher induced SVO values decrease the probability of stopping the game, reproducing the relationship between prosociality and cooperation observed in human behaviour. Together with the sensitivity of LLMs' elicited SVO to prompt and order effects, these results suggest that the usefulness of SVO for behavioural simulation does not depend on LLMs possessing stable social preferences, but rather on their ability to map social preferences onto corresponding strategic choices.