发表机构
University of Oxford; Department of Politics and International Relations, University of Oxford; Oxford Computational Political Science Group; Department of Political Science, Stanford University; LSE(牛津大学; 牛津大学政治与国际关系系; 牛津计算政治科学小组; 斯坦福大学政治学系; 伦敦政治经济学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究合成智能体能否复现联合实验捕捉的人类多维度偏好,通过复制已发表的联合研究,发现合成参与者的有效性依赖于研究主张,其无法充分替代人类样本。
AI 中文摘要
尽管人们越来越关注利用大语言模型(LLMs)提升调查实验的稳健性或降低数据收集成本,但它们在联合设计(政治学中一种日益流行的方法)中的有效性仍未得到充分探索。本文通过研究合成智能体能否复现联合设计旨在捕捉的多维度人类偏好模式,填补了这一空白。本文复制了已发表的联合研究,并在三个维度上比较了合成智能体生成的结果与原始人类数据:表征对应性、推理对应性和程序稳定性。我们的分析评估了选择分布的一致性,以及估计值的统计和实质相似性,结果在这些维度和所复制的研究之间并不一致。这意味着合成参与者的有效性应被视为依赖于研究主张且具有层级性。复现图表或获得强符号一致性是总体输出相似的证据,但不足以支持替代人类受访者。我们的结果表明,整个学科在将合成智能体视为人类样本的稳健替代品之前,必须先在各个层面上明确这一创新的边界。
英文摘要
Despite growing interest in using LLMs to add robustness or reduce data-collection costs in survey experiments, their efficacy in conjoint design---an increasingly popular method in political science---remains underexplored. This paper addresses that gap by investigating whether synthetic agents can reproduce the multi-dimensional human preference patterns that conjoint is designed to capture. It replicates published conjoint studies and compares the results generated by synthetic agents with original human data along three dimensions: representational correspondence, inferential correspondence, and procedural stability. Our analysis evaluates the alignment of choice distributions as well as the statistical and substantive similarity of estimates, and the results are uneven across these dimensions and studies replicated. This implies that the validity of synthetic participants should be considered claim-dependent and hierarchical. Reproducing a figure or obtaining strong sign agreement is evidence of similar aggregate outputs, but not enough to support replacing human respondents. Our results suggest that the discipline as a whole must first map this innovation's boundaries across various levels before considering synthetic agents a robust substitute for human samples.