大型语言模型大幅压缩福祉不平等,但在很大程度上保留其社会经济结构
Large Language Models Substantially Compress Well-Being Inequality but Largely Preserve Its Socioeconomic Structure
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究用六个LLM预测93,901名受访者的生活满意度,发现模型虽压缩总体离散度,但归一化后仍保留收入梯度等社会经济结构,表明压缩与结构保留可并存。
AI中文摘要:
使用大型语言模型(LLM)生成合成人群的研究一再表明,模型输出压缩了人类经验的多样性。这引发了人们对LLM生成的数据能否捕捉人群内部有意义差异的质疑。我们表明,这种压缩并不一定会抹去人类异质性的社会结构。利用世界价值观调查第七波中来自66个国家和地区的93,901名受访者,我们让六个LLM根据人口统计、社会经济和态度特征预测受访者的生活满意度。所有六个模型都大幅低估了生活满意度的整体离散程度。然而,在针对这些规模差异进行归一化后,它们在很大程度上再现了人类在福祉不平等方面的收入梯度:低收入群体仍然比高收入群体相对更加异质。这一模式对国家固定效应、国家等权加权、WVS调查权重以及观察到的人口构成具有稳健性,并在方向上延伸至就业、教育和感知控制。对于极端结果和国家特定的梯度,保真度较弱。这些结果表明,LLM保留的异质性数量以及这种异质性在社会群体间的分布方式是截然不同的属性。因此,LLM生成的人群可以在大幅压缩人类变异的同时,保留关于变异集中位置的有意义信息。
英文摘要:
Research using large language models (LLMs) to generate synthetic populations has repeatedly shown that model outputs compress the diversity of human experience. This has raised doubts about whether LLM-generated data can capture meaningful differences within populations. We show that such compression does not necessarily erase the social structure of human heterogeneity. Using 93,901 respondents from 66 countries and territories in Wave 7 of the World Values Survey, we ask six LLMs to predict respondents' life satisfaction from demographic, socioeconomic, and attitudinal profiles. All six models substantially understate the overall dispersion of life satisfaction. Yet after normalizing for these differences in scale, they largely reproduce the human income gradient in well-being inequality: lower-income groups remain relatively more heterogeneous than higher-income groups. The pattern is robust to country fixed effects, equal-country weighting, WVS survey weights, and observed demographic composition, and it extends directionally to employment, education, and perceived control. Fidelity is weaker for extreme outcomes and country-specific gradients. These results show that the amount of heterogeneity preserved by an LLM and the way that heterogeneity is distributed across social groups are distinct properties. LLM-generated populations can therefore substantially compress human variation while retaining meaningful information about where that variation is concentrated.