arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

看似合理但并非有效:对大型语言模型(LLMs)作为合成调查受访者的心理计量学审查

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

Mantas Lukauskas, Viktorija Šarkauskaitė

arXiv 2608.14606首次发表:更新:

发表机构

Hostinger; Kaunas University of Technology; AI Insight Lab(Hostinger; 考纳斯理工大学; AI Insight Lab)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过心理计量学分析发现,虽LLMs可复现人类调查关系的定性方向,但其心理计量学相似性不及高斯copula基线,且存在默许偏移、预测效度下降等问题,无法作为人类调查数据的直接替代品。

AI 中文摘要

大型语言模型(LLMs)正越来越多地被用作合成调查受访者,但现有评估仅关注单个层面上的答案是否看似合理。本文认为恰当的问题应是心理计量学层面的:LLMs是否保留了真实人类调查数据的联合分布、潜在结构、信度、中介路径和人口统计学效应?我们引入了一个立陶宛组织心理学数据集(包含263名员工;涉及Dunham变革态度量表、UWES-17、Koopmans IWPQ;共68个项目、12个子量表),并在五级角色披露阶梯下,对涵盖OpenAI、Anthropic、Google及12种开放权重模型家族的37个模型阵容,基于真实受访者特征进行了条件测试,同时开展了表现与推理工作量消融实验、反事实人口统计学交换(性别、角色、教育程度)、跨语言检查及逐字回忆记忆探测。所得的心理计量学相似性得分(PSS)以5个非LLM统计基线和保留的人-人上限为基准,结合受访者自助法置信区间和Tucker's phi的项目置换原假设进行评估。LLMs能复现人类心理计量学关系的定性方向,但高斯copula基线在样本驱动的PSS组件上优于所有LLMs;LLMs“群体”自身的相似性(平均LLM间PSS为0.73)高于与人类的相似性;记忆并非排行榜的驱动因素(回忆-PSS秩相关为0.00)。反事实交换显示,教育程度驱动的效应(平均|d|=0.56)远大于性别(0.12)和角色(0.18);37个模型中有8个的UWES的Tucker's phi落在置换原假设内。下游分析显示,所有LLMs均表现出强烈的默许偏移(+0.84 SD),基于合成数据训练的回归模型对保留人类数据的预测效度下降(平均R²为-0.18,对比人类数据的0.28),且模型在10条安慰剂中介路径中有3条编造了间接效应。LLMs样本并非人类调查数据的直接替代品。

英文摘要

Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.

Comments50 pages, 9 figures. Under review. Code and data will be released upon publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑