arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05828cs.AI

数据驱动的调查模拟人物角色:跨数据访问模式下的模拟对齐洞察

Data-Driven Personas for Survey Simulation: Insights into Simulation Alignment Across Data-Access Regimes

Dongryeol Lee, Weronika Łajewska, Leonardo Perelli, Saab Mansour

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究从异构匿名公共数据中诱导人物角色以模拟特定人口群体调查响应,发现域外数据因人口不匹配效果有限,而准确分配和目标域数据可显著提升模拟对齐。

中文摘要 AI 辅助

然而,许多现有的引导方法依赖于目标领域的人类数据进行微调或提示,这些数据收集成本高昂且引发隐私担忧。在本文中,我们研究了人口统计群体层面的调查模拟,其中从异构、匿名的公共行为数据中诱导出的人物角色,用于调节模拟特定人口统计群体个体反应的智能体。我们考察了是否可以从不同来源诱导出代表性人物角色,并分析了源数据的领域、规模和粒度如何影响调查模拟对齐。我们发现,从域外来源诱导的人物角色很少能优于仅基于基本人口统计信息进行条件化的模拟,这主要是由于人口不匹配。然而,当人物角色被准确分配到目标人口统计群体时,对齐显著改善。最后,从目标领域调查数据中诱导的人物角色,随着更多调查问题历史的可用而泛化得更好,这表明更丰富的行为证据能够实现更稳定的人物角色特质推断,从而转移到更好的未见问题模拟对齐。

英文摘要

Large language models (LLMs) offer new opportunities for public opinion research by enabling early prediction of survey responses, potentially reducing the cost and time of traditional surveys. However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heterogeneous, anonymized public behavioral data condition agents that simulate responses of individuals from specific demographic groups. We examine whether representative personas can be induced from diverse sources and analyze how the domain, scale, and granularity of the source data affect survey simulation alignment. We find that personas induced from out-of-domain sources rarely outperform simulations conditioned only on basic demographic information, largely due to population mismatch. However, when personas are accurately assigned to the target demographic groups, alignment improves substantially. Finally, personas induced from target-domain survey data generalize better as more survey question history becomes available, suggesting that richer behavioral evidence enables more stable persona trait inference that transfers to better unseen questions simulation alignment.

发表机构

  • Seoul National University(首尔大学)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑