何时可以信任你的合成用户?LLM消费者面板的诊断与修正
When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels
查看机构详情
- Recast(Recast公司)
- Dell(戴尔)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对LLM合成消费者面板的信任问题,提出分解偏差、诊断阈值与AIPW修正框架,经模拟及两个真实数据集验证,可有效识别并大幅降低偏差。
中文摘要 AI 辅助
大型语言模型日益被部署为合成消费者面板,有望比传统调查降低97%的成本。然而,总体验证指标掩盖了系统性失败:方差压缩、系数符号翻转、亚组误差膨胀10至30个百分点,以及全局修正反而加剧人口统计偏差。我们提供了一个正式框架,用于决定何时信任、修正或放弃LLM生成的消费者数据。该框架将合成面板偏差分解为协变量偏移和概念偏移,开发了具有可解释决策阈值的可测试诊断方法,并提供了一种仅需少量校准样本(n=50至300)的双重稳健AIPW估计器。我们在三个测试平台上进行了验证。在受控模拟中,决策规则达到了100%的准确率(180/180次重复)。在预先存在LLM失败的美国全国选举研究中,它正确标记了异质性概念偏移,并将朴素偏差减少了92.9%至99.6%。在Twin-2K-500消费者定价数据集(172,884对人工和GPT-4.1-mini响应)上,它正确地将全样本估计路由到“信任”,将亚组目标路由到“修正”,偏差减少了83%至94%。
英文摘要
Large language models are increasingly deployed as synthetic consumer panels, promising $97\%$ cost reductions over traditional surveys. Yet aggregate validation metrics conceal systematic failures: variance compression, coefficient sign-flips, subgroup error balloons of 10--30 percentage points, and global corrections that worsen demographic bias. We provide a formal framework for deciding when to trust, correct, or abandon LLM-generated consumer data. The framework decomposes synthetic-panel bias into covariate and concept shift, develops testable diagnostics with interpretable decision thresholds, and supplies a doubly robust AIPW estimator requiring only a small calibration sample ($n = 50$-$300$). We validate on three testbeds. In controlled simulations the decision rule achieves $100\%$ accuracy (180/180 replications). On the American National Election Study with pre-existing LLM failures, it correctly flags heterogeneous concept shift and reduces naive bias by $92.9-99.6\%$. On the Twin-2K-500 consumer pricing dataset (172,884 paired human and GPT-4.1-mini responses), it correctly routes full-sample estimation to Trust and subgroup targeting to Correct, with $83-94\%$ bias reduction.