发表机构
Seeburg Castle University; decision-context; Artificial Societies Ltd.; University of Oxford(塞堡城堡大学; 决策情境; 人工社会有限公司; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对合成调查验证错位问题,提出双要求框架:明确对应水平与亚组报告,操作化三个正义维度,并以电动汽车充电电价为例验证。
AI 中文摘要
工业界和学术界的研究人员使用由大型语言模型驱动的合成调查受访者作为人类样本的替代品。这些合成人群需要针对真实世界数据进行验证,因此研究人员通常通过与人调查的临时比较来处理。受行为科学中意图-行为差距的启发,我们认为,对于决策者委托合成研究以预测后果性行为的大多数应用场景,这些验证测试的对象是错误的。为解决这一问题,我们提出了一个包含两项要求的验证框架。首先,每个有效性声明必须说明其与人类数据的对应水平:样本是否能预测所代表人群的行为,验证针对四个诊断指标(位置、离散度、响应过程和结构)中的哪一个,以及验证是否与实验效应进行比较?其次,研究人员必须报告亚组的有效性声明,因为这些群体往往受后果性决策影响最大,而总体准确性掩盖了它们的错误表征。我们的验证框架将三个正义维度(分配性、程序性和承认性)操作化为可测量的量,并将人物内反事实实验定义为验证要求。然后,我们将该框架应用于电动汽车充电电价,最后提供一个报告清单,研究人员可用其提出令人信服的有效性声明。
英文摘要
Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and treats within-persona counterfactual experiments as a design that itself requires validation. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.
Comments19 pages, 1 figure