arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30030cs.CL

人工社会基准:合成研究的验证框架

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He

首次发表
浏览论文内容

中文总结 AI 辅助

提出人工社会基准框架,通过十一项效度测试和九个语言模型比较,评估合成人群是否支持预期研究分析,并生成记分卡以指导证据使用。

中文摘要 AI 辅助

合成调查可以再现平均答案,但同时可能歪曲人们之间的差异、答案之间的相互关系,或他们对条件变化的反应方式。我们引入了人工社会基准(Artificial Societies Benchmark),以帮助研究人员评估合成人群是否支持其预期的分析。该框架结合了内部效度、构念效度和外部效度方面的十一项测试,基于二十个人类数据源,并比较了九个语言模型。它将每项研究用途与其所需的证据联系起来,并测试当我们向受访者提供的信息发生变化时,结果如何改变。重要的是,在一个领域中的优异表现并不能确立其他领域的保真度。模型往往回答过于一致,压缩响应量表,并改变特征之间的关系,而更丰富的画像改善了一些模型的预测,却恶化了其他模型的预测。由此产生的记分卡帮助研究人员识别合成人群的哪些方面可以支持其分析,以及哪些方面需要进一步的人类证据。

英文摘要

A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.

发表机构

  • University of Oxford(牛津大学)
  • Artificial Societies(人工社会)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑