arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

营销研究中的合成数据:如何评估以及何时信任

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

Oded Netzer, Rajan Sambandam

arXiv 2609.13995首次发表:更新:

发表机构

Columbia Business School; TRC Insights(哥伦比亚商学院; TRC洞察公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出评估营销研究合成数据的方法,区分三种类型并引入遗忘问题场景,通过R²诊断筛选孪生数据,显著提升个体层面相关性并降低低质量问题比例。

AI 中文摘要

关于营销研究中合成数据的争论已经两极分化:一方面声称大型语言模型(LLM)使人类受访者变得过时,另一方面呼吁完全避免使用它们。我们认为这两种立场都掩盖了更有用的问题:不是合成受访者是否有效,而是何时有效。基于Brand、Israeli和Ngwe(2026)的研究,我们做出三项贡献。首先,我们区分了三种类型的合成数据(无根据的LLM响应、细分层面的人物画像和个体层面的数字孪生),并将每种类型映射到其可支持的决策。其次,我们开发了一个包含四类准确性度量的分类体系,并指出所报告的孪生准确性范围从接近完美到接近随机,很大程度上反映了测量对象的差异,而非方法质量的差异。即使提供给LLM的信息很少,聚合度量通常也表现良好,并且可能掩盖完全没有受访者层面区分的情况。第三,我们引入了遗忘问题问题,即某个问题在实地研究中被遗漏,作为基于孪生扩展现有数据的场景。我们提出了一种无需真实标签的事前可回答性诊断:用构建孪生的数据预测孪生输出的随机森林的R²。在一项具有全国代表性的调查(N=3,063)中的108个态度问题上,以R²高于0.7进行筛选,将孪生-人类个体层面相关性的平均值提高了15%,并将回答不佳问题的比例从25.9%降至4.3%。嵌入相似性和经验丰富研究者的判断提供了相关但较弱的筛选。

英文摘要

Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑