arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16395cs.CY

硅采样以国家层面的假设而非个体态度作答:来自欧洲社会调查的跨国证据

Silicon sampling answers with country-level assumptions, not individual attitudes: Cross-national evidence from the European Social Survey

  • London School of Economics and Political Science(伦敦政治经济学院)

机构由 AI 辅助整理,请以论文原文为准。

Chuyao Wang

AI总结:

本研究评估硅采样(用LLM模拟调查受访者)在跨国调查中的表现,发现其聚合还原度中等,国家标签起关键作用,但个体层面还原无效,仅适合探索性国家排名比较。

AI中文摘要:

硅采样使用大型语言模型(LLMs)模拟调查受访者。它能否还原跨国差异,以及为何如此,仍未得到解决。本研究以欧洲社会调查第11轮(30个国家,42个题项)为基准,在两种开放权重LLM下,采用第一人称和第三人称提示,并加入背景故事和回答格式实验,对其进行评估。总体还原度中等,且在各题项间分布不均。在包含三个变量的人口统计背景故事中加入国家名称,将模拟国家均值与观测国家均值之间的每题项中位相关系数从-0.03提升至0.52,而测试的更丰富画像未带来一致增益。受访者的国家标签充当一种国家层面的假设,受访者细节无法修正该假设。用文字标注回答量表的端点,可防止模型将国家排名倒置,因此回答格式决定了排名的方向。个体层面的还原度在所有条件下均可忽略不计,且不随跨国聚合还原度而变化。使用相邻国家的平均值(不借助LLM)在还原国家水平上比所有模型条件更准确,且排名效果与之相当。因此,硅采样可在题项层面验证并报告回答格式后,支持探索性的国家排名比较,但不支持个体或分布推断。

英文摘要:

Silicon sampling uses large language models (LLMs) to simulate survey respondents. Whether it recovers cross-national variation, and why, remains unresolved. This study evaluates it against European Social Survey Round 11 (30 countries, 42 items) with two open-weight LLMs under first- and third-person prompts, plus backstory and response-format experiments. Aggregate recovery is moderate and uneven across items. Adding the country name to a three-variable demographic backstory raises the median per-item correlation between simulated and observed country means from -0.03 to 0.52, and the richer profiles tested add no consistent gain. The respondent's country label acts as a country-level assumption that respondent detail does not revise. Naming the response-scale endpoints in words stops the model from ranking countries backwards, so the answer format sets the direction of the ranking. Individual-level recovery remains negligible in every condition and does not track aggregate recovery across countries. An average of neighboring countries, using no LLM, recovers country levels more accurately than every model condition and ranks them about as well. Silicon sampling can thus support exploratory country-ranking comparison after item-level validation and with the response format reported. It does not support individual or distributional inference.

补充信息

↑