arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24782cs.AI

价值对齐中的个性化、角色设定与预测

Personalization, Personas, and Forecasting in Value Alignment

James Wedgwood, Pratiksha Thaker, Neil Kale, Virginia Smith

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨LLM行为受人类身份影响的方式,通过WVS测试不同框架的互换性,在多语言多国家问题上评估多个模型,发现提示框架是文化对齐的关键因素,不同框架效果有别,对齐增益集中在部分价值维度,制度信任等问题仍难。

中文摘要 AI 辅助

大语言模型(LLM)的行为可能通过多种方式受到人类身份的影响:被要求适应用户、扮演特定人群或预测人们对价值负载问题的回答。我们使用世界价值观调查(WVS)来测试这些框架是否可互换。我们在13个语言-国家区域的101个源自WVS的问题上评估了GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash和Qwen3-235B,将仅语言基线与用户-国家、角色-国家和第三人称提示进行比较。在21008个模型响应行中,提示框架是文化对齐的一阶决定因素:国家线索常常大幅改变答案,但并非所有变化都朝着与人类响应分布匹配的方向。对于四个托管模型中的三个,第三人称预测产生了最强的方向对齐,而个性化和角色扮演则较弱或不太稳定。对齐增益集中在宗教信仰、性别角色和工作导向的物质价值等显著价值维度上,而制度信任和与民主相关的问题仍然困难。这些结果表明,提示框架在文化价值引出中不是一个表面的选择;它改变了模型行为和测量的对齐。

英文摘要

LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.

↑