发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过五种信息条件对比,发现行为历史比人物描述更能提升LLM合成人格对个体决策的预测准确性,尤其在个体层面显著增强预测能力。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用作代表调查受访者的合成人格。它们作为特定受访者替代品的有效性取决于它们能否重现个体的决策。我们考察了哪些信息有助于合成受访者预测每个个体后续的选择,使用了五种条件,逐步增加更丰富的信息:无个人信息、人口统计学特征、人格特质、认知得分,以及最后受访者早期的调查选择作为行为历史。我们使用了一个由845名美国成年人组成的两波面板数据,他们完成了14种行为偏差的测量(涵盖风险、时间偏好、过度自信和推理),因此每位受访者早期的回答提供了一个人类重测基准;在行为历史条件下,所有用于评分目标偏差的项目均被扣留。在总体水平上,每种条件下每位受访者的平均偏差数量接近人类平均值(7.1-8.1个偏差,而人类为7.1个)。这种总体相似性掩盖了方差上的差异:基于人物描述的条件仅恢复了人类个体间变异的53-67%,而添加行为历史则将其恢复到接近人类水平。在个体水平上,基于描述的人格仅达到人类重测反应中观察到的信息量(informedness)的7-12%,而添加行为历史则将其提升至28%。包含行为历史的条件在所有17个人口统计学群体中具有最高的估计信息量,而基于描述的条件对某些群体提供很少或没有信息。合成反应还表现出比人类反应更强的与教育和收入相关的差异。对于LLM合成人格,受访者过去的回答比对其是谁的描述更能增加个体层面的预测能力。
英文摘要
Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add progressively richer information: no personal information, demographics, personality traits, cognitive scores, and finally the respondent's earlier survey choices as behavioral history. We use a two-wave panel of 845 US adults who completed measures of 14 behavioral biases (spanning risk, time preferences, overconfidence, and reasoning), so each respondent's earlier answers provide a human test-retest benchmark; in the behavioral-history condition, all items that score the target bias are withheld. At the population level, the average number of biases per respondent in every condition is close to the human average (7.1-8.1 biases, against 7.1 for humans). This aggregate similarity masks differences in variance: persona descriptions recover only 53-67% of human between-person variation, whereas adding behavioral history restores it to approximately the human level. At the individual level, description-based personas achieve only 7-12% of the informedness observed in human test-retest responses, while adding behavioral history raises this to 28%. The condition including behavioral history has the highest estimated informedness in all 17 demographic groups, whereas description-based conditions provide little or no information for some groups. Synthetic responses also exhibit stronger education- and income-related differences than human responses. For LLM synthetic personas, a respondent's past answers add more to individual-level prediction than a description of who they are.