模拟社会的局限性:后训练与调查微调如何抹去跨文化差异
The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance
浏览论文内容
中文总结 AI 辅助
本文揭示后训练与调查微调导致LLM模拟社会时出现“共识崩溃”,压缩意见多样性,以点准确率衡量会误判模拟器,当前训练牺牲多样性换取共识。
中文摘要 AI 辅助
使用大型语言模型(LLM)模拟多样化的人类群体,有潜力改变计算社会科学的许多方面,然而许多评估只关注平均响应,而非真实群体内意见的分布。在此,我们开发了一个诊断框架,该框架在来自世界价值观调查(WVS)的10,000个受访者-问题对上,同时测量点准确性和离散度保持率(预测标准差与人类标准差之比,记为$\dr$),这些数据覆盖十二个国家和六大洲。我们评估了十一个零样本语言模型和五个在WVS数据上使用SFT、DPO和GRPO进行微调的变体。我们识别出一种我们称之为“共识崩溃”的失败模式,即对齐训练将输出压缩为每个群体的一种刻板印象。沿着从Llama 3.1 70B基础模型到Tulu 3检查点的后训练轨迹,第一阶段,即监督指令微调,在准确率提升极小的情况下($\dr$从1.22降至0.59;准确率提升+0.9个百分点)移除了一半的离散度,后续阶段并未恢复该离散度,并且WEIRD国家与非WEIRD国家之间出现了差距,而调查微调在追求更高点准确率的同时加深了这一差距。最准确的模型(在WVS上微调的Tulu 3 70B-DPO,准确率57.9%)整体保留了人类离散度的一半($\dr = 0.50$),对尼日利亚仅保留了11%,而WEIRD国家则为0.70至0.87。将采样温度提高到1.0,对于两个微调的DPO模型,其与人类分布的Wasserstein-1距离($\wone$)保持不变,而在Qwen 3.5 9B上使用GRPO,无论是在准确率奖励还是分布形状奖励下,均未恢复离散度。将对齐模型与未对齐的先验混合,在保留的测试集上将$\dr$从0.51提高到0.62,但尼日利亚仍为0.36。因此,仅凭点准确率会误判这些模拟器,而当前的后训练以多样性换取共识。
英文摘要
Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\%) keeps half the human spread overall ($\dr = 0.50$) and 11\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.
发表机构
- Georgetown University(乔治城大学)
机构由 AI 辅助整理,请以论文原文为准。