发表机构
University of Georgia; MIT Sloan School of Management; Boston University(佐治亚大学; 麻省理工学院斯隆管理学院; 波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究验证LLM作为人类替代物的前提,发现其仅能近似项目均值,无法捕捉受访者特异性偏差,且更丰富的角色数据等无法缩小差距,提出四项实证检验方法。
AI 中文摘要
大语言模型(LLMs)越来越多地被用作人类替代物,其前提通常是更丰富的角色(persona)数据能使它们成为特定个体的替代品或探索工具。我们在四个覆盖超过40万名参与者、6000多项调查项目及实验结果的数据集上验证了这一前提。LLMs在总体层面表现良好:它们的平均响应与人类对相同项目的平均响应高度一致。但这一成功很大程度上反映了对每个项目人类平均响应的预测。一旦去除每个项目的人类均值,LLM预测仅能解释剩余的受访者特异性变异的3.05%,远低于53.6%的人类重测基准。更丰富的角色数据、模型变体及微调均无法缩小这一差距。方差分析显示,去除项目均值后,可靠的剩余信号为项目-受访者交互项,它捕捉了受访者在特定项目上偏离均值的程度,其规模约为稳定个体效应的8.9倍。角色数据编码了受访者,但未编码这种项目特异性偏差。LLM响应还会压缩人类响应分布,表现为更窄的分布范围、更少的响应类别及扭曲的分布形态。我们将此模式称为项目均值替代。当前的LLM替代物可近似项目均值,但无法提供替代个体人类所需的分布或受访者特异性偏差。我们为基于LLM的人类替代物主张提出了四项实证检验方法。
英文摘要
LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.
Comments37 pages, 6 figures, 14 tables. Includes Supporting Information (main text pages 1-9; SI Appendix pages 1-28)