发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型观点多样性,通过析因实验分离输入条件设定和交互架构两维度,在7个模型上评估,发现更多角色细节非单调增多样性,不同架构探索不同观点区,组合架构覆盖更广,低成本措施效果差,揭示多样性受干预结构形式和组合影响。
AI 中文摘要
大语言模型越来越多地用于模拟开放式任务中的不同人类观点。然而,其输出存在系统性的观点同质化。从业者探索了多种干预措施,但评估方式零散。为更科学地理解大语言模型输出的多样性,设计析因实验,分离输入条件设定(通过角色深度实现)和交互架构两个主要干预维度。在7个模型上针对100个真实用户开放式问题评估所有条件,用多种互补指标衡量多样性。研究结果挑战了一些常见假设,如更多角色细节不一定单调增加多样性,不同架构探索不同观点区域,组合多种架构覆盖更广,低成本替代措施效果可忽略不计。表明多样性并非单一维度扩展的产物,对干预的结构形式和组合高度敏感。
英文摘要
Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.