arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模拟还是估计?基础语言模型与后训练语言模型在观点模拟中的差异化优势

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson

arXiv 2608.03044首次发表:更新:

AI 中文总结

该研究针对大型语言模型用于观点模拟的矛盾结果,区分了模拟与估计任务,发现基础模型更适合模拟、后训练模型更适合估计,为相关模型选择提供了依据。

AI 中文摘要

大型语言模型正越来越多地被用于模拟人类观点,但已有研究报告了相互矛盾的结果:部分研究发现其与人类调查数据具有良好的对齐性,而另一些研究则发现存在角色崩溃和人口统计敏感性较弱的问题。本文表明,这种冲突很大程度上源于混淆了两项不同的任务。我们将第一项任务称为模拟(emulation),即模型生成的个体响应汇总为人口分布;第二项任务称为估计(estimation),即模型直接预测人口分布。我们在Pew美国趋势小组(Pew American Trends Panel)上评估了6组匹配的基础模型(base models)和后训练模型(post-trained models),发现基础模型是更强的模拟器:它们生成的响应分布更接近人类真实值,且能更好地保留人口统计结构;后训练模型则是更强的估计器,在被要求直接预测时能产生更准确的分布预测。我们提出,针对人类模拟的模型选择应根据任务是否需要生成文本或预测分布来指导。

英文摘要

Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We propose that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are generally the stronger estimators, producing more accurate distributional predictions when asked directly. We argue that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.

CommentsEMNLP Findings 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑