发表机构
Columbia University; University of Pennsylvania(哥伦比亚大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模受试者间实验,证明AI主持访谈在深度和主题覆盖上媲美人类主持,且成本效益更高,但情感参与度较低;其数字孪生在定量预测上不优于静态访谈,误差源于思维风格差异与分布外问题。
AI 中文摘要
AI主持的访谈正作为一种可扩展的市场研究方法出现,用于生成消费者洞察并构建消费者“数字孪生”。然而,目前尚不清楚它们是否能与人类主持的访谈相媲美,或能否优于更简单的静态数据收集方法。在一项预先注册的、受试者间设计的研究(N = 317)中,我们与三家行业合作伙伴合作,比较了AI主持(N = 139)、人类主持(N = 24)和静态访谈(N = 154)。AI主持在深度上与人类主持相当,覆盖了更多主题,并且在预算固定的情况下,比人类主持或静态访谈能显著恢复更多的客户需求。然而,当参与者与真人交谈时,他们的声音听起来情感参与度更高。随后,我们使用访谈数据创建数字孪生,并根据参与者自己对六个真实世界营销刺激的保留回答来评估每个孪生。我们发现,由AI主持的访谈创建的数字孪生在预测消费者反应方面优于仅基于人口统计特征的画像。然而,与静态访谈相比,AI主持带来的额外丰富性并未转化为更好的定量预测。通过分析人类与其孪生生成的无限制思考,我们发现预测误差既与孪生和人类之间(自我报告的)思维风格的差异有关,也与训练数据和验证数据之间的差距(即提出分布外的问题)有关。
英文摘要
AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer "digital twins." Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).