arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解基于LLM的人类行为模拟中的可靠性

Understanding Reliability in LLM-based Human Behavior Simulation

Pei Wang, Lei Wang, Yuanzi Li, Xu Chen

arXiv 2609.25066首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出ReliMap框架,将基于LLM的人类行为模拟分解为三个层面,发现画像条件化可减少分布偏差但收益递减,且个体与群体层面可靠性可能背离,需协调改进所有层面才能实现可靠模拟。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用于模拟人类调查响应和行为反应,然而不可靠的模拟可能会误导社会科学结论。然而,现有评估侧重于端到端得分,使得模拟过程的不同方面如何相互作用以决定可靠性仍不清楚。我们提出ReliMap,它将基于LLM的人类行为模拟分解为三个结构化层,并在三个配置维度(模型能力、画像完整性和群体覆盖度)上,在个体层面(R1)和群体层面(R2)评估可靠性。通过在四个模拟任务和十一个LLM上的实验,我们发现,在没有画像条件化的情况下,所有模型都表现出显著的分布偏差。画像条件化以递减的收益减少了这种偏差。更大的模型受益更多,且属性信息量比数量更重要。关键的是,R1的增益不能可靠地转移到R2——个体和群体层面的可靠性可能朝相反方向移动。在群体层面,增加覆盖度减少了方差但未减少系统性偏差,R2在约50-100个个体时趋于稳定。这些发现强调,可靠的模拟不能通过单独优化任何单一层面来实现,而需要跨所有三个层面的协调改进。

英文摘要

Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completeness, and population coverage. Through experiments across four simulation tasks and eleven LLMs, we find that all models exhibit substantial distributional bias without profile conditioning. Profile conditioning reduces this bias with diminishing returns. Larger models benefit more, and attribute informativeness matters more than quantity. Critically, R1 gains do not reliably transfer to R2--individual and population-level reliability can move in opposite directions. At the population layer, increasing coverage reduces variance but not systematic bias, with R2 stabilizing at around 50-100 individuals. These findings highlight that reliable simulation cannot be achieved by optimizing any single layer in isolation, but requires coordinated improvement across all three.

Journal refThe 2026 Conference on Empirical Methods in Natural Language Processing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑