发表机构
Indiana University Bloomington; Massachusetts Institute of Technology; Root Dynamix LLC(印第安纳大学布卢明顿分校; 麻省理工学院; Root Dynamix有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估了九种开放权重大语言模型在三个选择领域生成的偏好分布,发现模型间存在显著不一致,且模型选择比提示措辞影响更大,挑战了LLM可替代受访者的假设。
AI 中文摘要
大语言模型(LLMs)越来越多地被用作概率生成器,用于在真实世界数据不可用的场景中进行模拟、合成数据生成和决策支持。然而,它们产生的分布的结构和可靠性仍未得到充分研究。在此,我们系统地分析了LLM生成的关于航空旅行、餐厅和消费品的偏好分布。令人鼓舞的是,我们分析中考虑的所有模型都表现出自我一致性,最可能的结果在重复采样下迅速稳定。与此同时,我们观察到模型家族和规模之间存在显著的不一致,即使在其最可能的结果中也几乎没有共识。这些模式在九个开放权重模型、三个选择领域中保持一致,并在温度变化、贪婪解码以及提示和顺序扰动下表现出鲁棒性。我们的发现表明,结果受模型选择的影响大于受提示措辞的影响,这挑战了常见假设,即足够强大的LLM在用作调查受访者的替代品时会产生相似的偏好分布。
英文摘要
Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, all models considered in our analysis exhibit self-coherence, with the most probable outcomes stabilizing rapidly under repeated sampling. At the same time, we observe substantial discordance across both model families and scales, with little consensus even among their most probable outcomes. These patterns hold across nine open-weight models, three choice domains, and show robustness under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that outcomes are influenced more by the choice of model than by the wording of the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.