arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

经验基础提升了在干扰期间模拟人类行为的大语言模型智能体的现实性

Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions

Chen Xia, Zexi Kuang, Yuqing Hu

arXiv 2607.17437首次发表:更新:

AI 中文总结

研究探讨经验基础对提升LLM智能体模拟人类行为现实性的作用,开发基于经验的框架,将多项调查数据嵌入其中,通过实验验证该框架能显著提升模拟效果,使LLM智能体成为更具统计可信度的人口行为模拟器。

AI 中文摘要

大语言模型(LLM)智能体为模拟人类行为提供了一种生成方法,这在灾害和基础设施中断规划中是常见挑战。但这种生成能力存在有效性问题。本文评估经验基础是否能提升LLM智能体模拟的统计现实性。具体开发了一个基于经验的LLM智能体框架,将多项调查数据嵌入智能体初始化等环节。通过2024年7月费城热浪期间的独立家庭调查验证,结果表明该框架能显著提升模拟效果,揭示了建模人类适应干扰时的差距。

英文摘要

Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. However, this generative capacity creates a validity problem: individually plausible agent reasoning may fail to reproduce empirical population behavior. We evaluate whether empirical grounding improves the statistical realism of LLM-agent simulations during disruptions. Specifically, we develop an empirically grounded LLM-agent framework that embeds demographic profiles from the American Community Survey, baseline routines from the American Time Use Survey, and urban spatial context into agent initialization, memory, decision prompts, and activity execution. An independent household survey conducted during the July 2024 Philadelphia heatwave is reserved as an external validation benchmark. Compared with an ungrounded LLM-agent baseline, the grounded model improved reconstruction of normal daily routines, increasing mean correlation with empirical activity profiles from 0.528 to 0.912 and reducing mean squared error from 0.066 to 0.008. Under heatwave conditions, the grounded model better reproduced survey-derived activity profiles, increasing mean correlation from 0.349 to 0.836 and reducing mean squared error from 0.098 to 0.012. The grounded model captured 46.4% of observed heatwave response amplitude, compared with 20.6% for the ungrounded baseline. These findings show that empirical grounding can make LLM agents more statistically credible simulators of population behavior while revealing remaining gaps in modeling human adaptation during disruptions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑