发表机构
University of Science and Technology of China; Fudan University(中国科学技术大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有方法无法模拟受访者在完整问卷中连贯回答的问题,提出FR-LLM,通过边际模型与自回归模型结合MCJP投影生成连贯问卷,在真实调查数据上提升多问题准确性并实现最高商业利润。
AI 中文摘要
大语言模型越来越多地被用于模拟社会调查中的回答分布。先前的工作已经实现了对单个问题的准确总体层面模拟。然而,真实问卷会向每位受访者提出一系列相关问题。一个模拟的受访者应在整个问卷中展现出连贯的偏好,而不仅仅是针对孤立条目的准确分布。现有的单条目方法无法准确再现同一个人回答完整调查的方式。我们提出了FullRespondent-LLM(FR-LLM),该方法微调了两个专门的LLM:一个用于每个条目回答分布的边际模型,以及一个用于跨答案依赖关系的受访者层面自回归模型。随后,边际约束联合投影(MCJP)将自回归联合分布投影到满足由第一个模型学习的条目层面边际的集合上。这产生了具有真实跨条目关系的完整问卷,同时保持了强大的条目层面准确性。在两个真实世界的社会调查数据集上,FR-LLM更准确地再现了多问题回答模式,保持了具有竞争力的单条目准确性,并更好地泛化到未见过的群体和问题。在一个小型商业调查数据集中,我们使用模拟回答来做出定价和库存决策;FR-LLM实现了最高的实际利润。
英文摘要
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
Comments20 pages, 6 figures