大型语言模型在模拟乳腺癌筛查干预后调查反应方面的表现如何?
How Well Do LLMs Simulate Survey Responses Following a Breast Cancer Screening Intervention?
浏览论文内容
中文总结 AI 辅助
本研究利用LLM智能体模拟乳腺癌筛查干预后的调查反应,发现基于档案的智能体优于基线但不及真实样本,且对特定人群和文化构念的模拟存在局限。
中文摘要 AI 辅助
收集调查数据既费时又受隐私限制。大型语言模型(LLMs)在预测性社会模拟方面已显示出潜力,但尚不清楚它们能否在医疗干预前后复制人群层面的反应分布。利用来自4125名年龄在35-59岁女性的信息,我们评估了仅依据干预前档案信息构建的智能体能否再现干预后的反应分布。我们使用Gemma 4 E4B和Qwen3.5 9B创建了LLM智能体组(每组n=50),条件范围从零样本提示到包含聚合或个体层面人口统计特征及干预前问卷反应的增强型智能体档案。我们通过总变差距离(TVD)和归一化Wasserstein距离(NWD)比较了预测与观察到的反应分布。在两种LLM中,基于档案的智能体相比零样本和随机基线提高了分布准确性。然而,直接抽样50名真实参与者仍然更准确。在55-59岁参与者及居住于私人房产的参与者中,预测误差也更高。误差还随问题主题和LLM模型而变化,其中癌症宿命论和干预后对遗传学态度的误差最高。敏感性分析显示,性能受提示模板变化和温度超参数影响。我们的结果表明,基于LLM的智能体在模拟干预行为反应方面具有潜力。然而,包含除人口统计信息之外额外信息的档案并未持续优于简单档案。某些文化构念和人群群体在所评估的LLM模型中仍未被充分代表。未来工作可能包括构建基于行为学并经过本地验证的虚拟人群。
英文摘要
Collecting survey data is laborious and limited by privacy constraints. Large language models (LLMs) have shown promise as predictive social simulations. It is unclear whether they can replicate population-level response distributions before and after a healthcare intervention. Using information derived from 4125 women aged 35-59 years, we evaluate whether agents informed solely by pre-intervention profile information can reproduce post-intervention response distributions. Groups of LLM agents (n=50) were created with Gemma 4 E4B and Qwen3.5 9B; conditions ranged from zero-shot prompting to agent profiles enriched with aggregate or individual-level demographic characteristics and pre-intervention questionnaire responses. We compared predicted and observed response distributions with Total Variation Distance (TVD) and Normalized Wasserstein Distance (NWD). Across both LLMs, profile-based agents improved distributional accuracy relative to zero-shot and random baselines. Nevertheless, direct sampling of 50 real participants remained more accurate. Prediction errors were also higher among participants aged 55-59 years and those living in private property. Errors also varied by question theme and LLM model, with the highest errors observed for cancer fatalism and post intervention attitudes toward genetics. Sensitivity analyses showed that performance was influenced by prompt template changes and temperature hyperparameter. Our results show the potential of LLM-based agents to model behavioral responses to interventions in silico. However, profiles containing additional information beyond demographics did not consistently outperform simpler ones. Certain cultural constructs and population groups also remain inadequately represented by the LLM models evaluated. Future work may include building behaviorally grounded and locally validated virtual populations.
发表机构
- Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR)(新加坡基因组研究所,科学技术研究局)
- Department of Statistics and Data Science, National University of Singapore(新加坡国立大学统计与数据科学系)
- Institute for Human Development and Potential (IHDP), Agency for Science, Technology and Research (A*STAR)(人类发展与潜力研究所,科学技术研究局)
- Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学李光耀公共卫生学院)
- Department of Surgery, Yong Loo Lin School of Medicine, National University of Singapore and National University Health System(新加坡国立大学杨鲁林医学院外科,新加坡国立卫生系统)
- Department of Surgery, National University Hospital and National University Health System(新加坡国立医院外科,新加坡国立卫生系统)
- National Cancer Centre Singapore (NCCS), Singapore Health Services (SingHealth)(新加坡国家癌症中心,新加坡卫生服务)
机构由 AI 辅助整理,请以论文原文为准。