arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10280cs.CY

总模拟调查误差:设计和诊断来自大语言模型的调查响应

Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models

发表机构曼海姆大学 · GESIS-莱布尼茨社会科学研究所 · 格拉茨大学
另 2 家 · 查看机构详情
  • University of Mannheim(曼海姆大学)
  • GESIS - Leibniz Institute for the Social Sciences(GESIS-莱布尼茨社会科学研究所)
  • University of Graz(格拉茨大学)
  • University of Duisburg-Essen(杜伊斯堡-埃森大学)
  • Complexity Science Hub(复杂性科学中心)

机构由 AI 辅助整理,请以论文原文为准。

Indira Sen, Georg Ahnert, Leah von der Heyde, Jana Lasser, Bernd Weiß, Markus Strohmaier

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型生成调查响应中的误差,提出总模拟调查误差框架,系统识别和反思设计生命周期中的概念错误与偏见。

中文摘要 AI 辅助

大语言模型(LLMs)在大量人类生成数据上训练后,可能编码了这些人类的态度和行为。因此,LLMs在模拟人类模式方面展现出潜力,使其能够在各种情境中用于模拟人类。其中一个情境是将LLMs用作“硅样本”,即作为回答调查问题以确立公众舆论、设计政策或用作(社会)科学数据的人类代理。然而,关于社会偏见、泛化性和技术局限性的几个关键问题仍然存在,并且由于模拟设计者面临广阔的设计空间而变得更加复杂。多宇宙分析可能有助于我们理解不同设计选择的影响,然而,我们缺乏对LLM生成调查的设计空间以及这些决策如何与固有的LLM局限性相互作用的系统性理解。因此,我们如何系统地识别、追踪和记录LLM生成调查响应中的局限性?基于定量社会科学中的传统,特别是调查方法和测量理论,我们调查了对LLM生成调查响应有效性的威胁。为此,我们设计了一个框架,枚举了在调查模拟生命周期不同阶段可能发生的概念性错误和系统性偏见。我们的框架,称为总模拟调查误差(TS2E)框架,提供了对LLM生成调查数据的统一且端到端的视角。该框架通过理论和实证案例研究进行说明,使调查模拟设计者能够系统地识别和反思LLM生成调查中的错误。

英文摘要

Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theoretical and empirical case study, enables survey simulation designers to systematically identify and reflect on errors in LLM-generated surveys.

补充信息

↑