arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26348cs.CLcs.AIcs.CYcs.HC

当合成用户失效时:LLM模拟人类调查回复的跨领域基准测试

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

Zihan Chen, Di Zhu, Lei Nico Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建跨领域基准与评估框架,发现LLM模拟人类调查回复存在个体表现逊于基线、过度确定人口统计特征等失效问题,且无法通过更大模型纠正,会误导细分目标决策。

中文摘要 AI 辅助

大型语言模型(LLM)正日益被用作合成用户,替代人类受访者,其模拟答案为产品、政策和市场决策提供依据。本研究探究这种替代何时有效、何时失效,并将答案打包为智能合成用户系统的评估框架。采用单一方案,覆盖两个模型家族的4个模型(规模从8B到前沿能力级别),应用于两个独立的真实人类响应数据领域:美国一般社会态度(综合社会调查General Social Survey)和跨文化价值观(世界价值观调查World Values Survey)。每个模型均与保留人类数据上拟合的一组非LLM基线进行基准测试。在人口统计提示和测试的调查模拟方案下,两种失效情况在两个领域、所有4个模型及两个家族中均重复出现:其一,个体层面,没有LLM优于最强基线;在跨文化价值观方面,所有模型的表现远低于基线,且该差距在距离感知和适当评分下依然存在。其二,模型系统性地过度确定人口统计特征,将身份视为态度的预测性远高于真实人群中的情况,这种失真存在于几乎所有问题-组组合中,且对编码不变量测量具有鲁棒性。更大、能力更强的模型也无法纠正这两种失效。决策影响分析显示其实际重要性:在细分目标任务中,模型将细分群体间差距放大2至4倍,会在一半美国案例和大多数跨文化案例中引导团队选择错误细分群体,并制造真实人群中不存在的细分群体划分。本研究将跨领域基准和评估框架按需提供,以便团队提前确定合成用户证据何时可安全用于决策支持,何时不可。

英文摘要

Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We publicly release the cross-domain benchmark and validation framework, including all code and data, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.

发表机构

  • Stevens Institute of Technology(史蒂文斯理工学院)
  • University of Massachusetts Boston(马萨诸塞大学波士顿分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑