arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型作为隐含的社会学模型:从社会人口统计特征重建投票行为

Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles

Roman Neruda, Martin Bakoš, Josef Šlerka, Vít Tuček, Petra Vidnerová, Gabriela Kadlecová

arXiv 2608.15871首次发表:更新:

发表机构

Institute of Computer Science, Czech Academy of Sciences; Faculty of Arts, Charles University(捷克科学院计算机研究所; 查理大学艺术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出将大型语言模型作为隐含社会学模型的方法论框架,以2021年捷克议会选举为案例,证实其可从社会人口统计特征重建投票行为,为计算社会科学提供新探索工具并明确其局限。

AI 中文摘要

在大规模互联网语料库上训练的大型语言模型(LLMs)编码了关于社会身份、态度和政治行为的广泛统计规律。本文引入并评估了一种方法框架,该框架利用这些潜在表征从个体层面的社会人口统计特征重建总体投票行为。我们通过将LLMs以人口统计描述为条件、引出概率性投票率和政党偏好,并通过软投票程序聚合个体输出,将LLMs作为隐含的社会学模型进行操作化。以2021年捷克议会选举为验证案例,我们证明当代LLMs能以较低的平均绝对误差重现官方选举结果、恢复已知的政治集团结构,并与独立确定的社会人口统计梯度相一致。本研究的贡献是方法论层面而非预测层面的:我们展示了如何系统地将LLMs作为社会现实的压缩表征进行探究,为计算社会科学提供了一种新颖的探索工具,同时明确界定了其认知和伦理局限。

英文摘要

Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we show how LLMs can be systematically interrogated as compressed representations of social reality, offering a novel exploratory instrument for computational social science while clearly delineating its epistemic and ethical limits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑