LLM数字孪生何时能减少人类测量?从行为保真度到统计可替代性
When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability
浏览论文内容
中文总结 AI 辅助
本研究提出统计可替代性标准,通过混合受试者与预测驱动推断框架评估LLM数字孪生减少人类测量的能力,发现行为保真度不等于统计可替代性,应基于推断有效性而非结果复现来评估AI证据。
中文摘要 AI 辅助
基于LLM的数字孪生有望通过生成个体特定响应来减少重复的人类数据收集,然而现有评估几乎没有提供证据表明它们能否在保持有效推断的同时减少人类测量。为解决这一问题,我们引入了统计可替代性(statistical substitutability),这是一种推断性标准,用于评估孪生预测在特定估计目标下能在多大程度上减少人类测量,同时保持有效推断。我们开发了一个基于混合受试者(mixed-subject)和预测驱动推断(prediction-powered inference)的框架,从四个维度评估统计可替代性:总体保真度、配对受访者层面的信号、有限样本下人类标签的恢复能力以及跨人群的稳定性。在涵盖行为实验、多种模型和替代性受访者表征的两项实证评估中,我们发现数字孪生能够再现人类平均效应,但几乎不提供关于哪些个体偏离这些平均值的信息。较新的模型和更丰富的受访者信息改善了某些维度的性能,但并未可靠地转化为人类数据节省。人类校准可以降低总体预测误差,但有限的标注样本往往无法产生稳定的精度提升。重要的是,这些发现表明行为保真度既不是统计可替代性的充分条件,也不是其必要条件。更广泛地说,它们表明AI生成的证据应基于其支持有效科学推断的能力来评估,而非仅仅基于其再现人类结果的能力。因此,数字孪生用于确认性用途时,应通过它们是否减少关于人类数量的不确定性来判断,而不仅仅是通过它们是否再现人类均值、分布或效应来判断。
英文摘要
LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.
发表机构
- University at Buffalo(布法罗大学)
机构由 AI 辅助整理,请以论文原文为准。