arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32638cs.AI

开放权重大语言模型(LLMs)能否模拟人类调查总体?一项跨工具校准研究

Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study

Grandee Lee, Wang Yue

AI总结:

本研究通过跨工具校准任务评估开放权重LLM模拟人类调查群体的能力,发现未调优模型即可复现真实人类跨工具相关性(r≈0.70-0.73),但性能不随版本单调提升,需版本特定的验证。

AI中文摘要:

大语言模型(LLMs)越来越多地被用于生成合成调查受访者和真实个体的数字孪生,但其输出是否保留了真实人类的统计结构(而非仅仅是表面上的合理性)仍未得到解决,且现有的大多数证据来自专有模型而非开放权重模型。我们在一个跨工具校准任务上评估了三个开放权重LLM家族:将人物画像条件化于真实受访者对一种心理测量工具的逐字回答,并在第二种受构念距离控制的工具上对其进行测量,与一个包含2,058人的真实人类小组进行对照检验。在139对网格中,模拟的跨工具相关性在每个模型中都追踪到了真实人类相关性,相关系数r = 0.70 - 0.73,这主要归因于正确的符号方向而非精确的数值大小,并且集中在构念距离适中的配对中。这种量级的相关性,仅通过基于个体层面调查数据条件化的未调优开放权重模型即可获得,对于基于LLM的行为模拟和数字孪生应用而言是一个实质性的鼓舞人心的结果:特定的模型家族和版本已经无需任何微调即可再现真实人类跨工具结构中有意义的一部分。然而,这种能力并不会随着模型版本的更新而单调提升:在一个匹配的小组上,所测试的三个Llama版本中最新的版本在三个主要指标中的两个上表现最差,因此要在实践中实现其潜力,需要针对特定版本且考虑距离的验证,而非一次性的基准测试。

英文摘要:

Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary models rather than open-weight ones. We evaluate three open-weight LLM families on a cross-instrument calibration task: conditioning personas on real respondents' verbatim answers to one psychometric instrument and measuring them on a second, construct-distance-controlled instrument, checked against a 2,058-person human panel. Across a 139-pair grid, the simulated cross-instrument correlation tracks the real human correlation at r = 0.70 - 0.73 in every model, driven mainly by correct sign rather than precise magnitude and concentrated in pairs of moderate construct distance. A correlation of this magnitude, obtained from untuned open-weight models conditioned only on individual-level survey data, is a substantively encouraging result for LLM-based behavioral simulation and digital-twin applications: specific model families and releases already reproduce a meaningful share of real human cross-instrument structure without any fine-tuning. This capability does not, however, improve monotonically across model releases: on a matched panel, the newest of three tested Llama releases performs worst on two of three headline metrics, so realizing its promise in practice requires release-specific, distance-aware verification rather than a one-time benchmark.

↑