AI 中文总结
研究对比GPT生成的调查问卷和既定调查基线在三个社会领域的情况,发现GPT生成的问卷与人工设计工具能捕捉相同主要态度划分,但有差异,适合探索性和大规模分析,可补充专家设计工具。
AI 中文摘要
理解人类信仰和社会态度通常依赖精心设计的调查工具。近期研究表明大语言模型(LLMs)可通过大规模生成调查问卷使部分流程自动化,这引发了此类工具与基于文献、人工设计的调查问卷可比性的问题。我们对GPT生成的调查问卷与既定调查基线在气候变化、移民、多元化、公平与包容这三个社会领域进行了对照实证比较。GPT生成的调查问卷采用固定提示框架,对信念、认知和行为采用3x3结构,而人工基线则从经过验证的工具中选取以匹配调查长度和结构覆盖范围。我们收集了美国参与者对两种调查问卷的回答以进行直接的主体内比较。我们分析了回答分布、聚类行为以及与自我认定立场的一致性方面的差异。结果表明,GPT生成的调查问卷与人工设计的工具捕捉到相同的主要态度划分,但在信念结构分辨率和群体区分方面存在差异。这些发现表明,大语言模型生成的调查问卷适用于探索性和大规模分析,可用于补充专家设计的工具。
英文摘要
Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate parts of this process by generating surveys at scale, raising questions about the comparability of such instruments to literature-grounded, human-designed surveys. We present a controlled empirical comparison between GPT-generated surveys and established survey baselines across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). GPT-generated surveys were produced using a fixed prompting framework enforcing a 3x3 structure over beliefs, perceptions, and behaviors, while human baselines were assembled from validated instruments to match survey length and construct coverage. We collected responses from U.S.-based participants, who completed both survey types, allowing direct within-subject comparison. We analyze differences in response distributions, clustering behavior, and alignment with self-identified stances. Our results show that GPT-generated surveys capture the same dominant attitudinal divisions as human-designed instruments, while exhibiting differences in the resolution of belief structure and group separation. These findings suggest that LLM-generated surveys are suited for exploratory and large-scale analyses, and can be used to complement expert-designed instruments.
DOI:10.63317/4juoym3quhm7 10.63317/4juoym3quhm7