发表机构
Johns Hopkins University; Listen Labs(约翰斯·霍普金斯大学; Listen Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发了经实证验证的AI面试官模拟环境InterviewPlayground,通过与15项含450名人类参与者的真实定性研究对比,证实其评估结果与人类研究表现的平均皮尔逊相关系数达0.86,可有效评估AI面试官性能。
AI 中文摘要
目前,人们正在开发越来越多的AI面试官,用于市场研究、民意调查、偏好 elicitation(偏好 elicitation 为专业术语,保留)和社会科学研究等场景中引出开放式回答。然而,评估AI面试官颇具挑战性,因为它们在长时间的多轮交互中运行,必须适应参与者的行为。为满足这一需求,我们开发了InterviewPlayground,这是一个基于社会理论构建的模拟研究参与者行为、用于评估AI面试官的模拟环境。在InterviewPlayground中进行的模拟研究会生成InterviewReportCard,该报告使用一套经过验证的指标评估AI面试官的性能。为测试我们基于模拟的评估是否能预测与人类参与者的表现,我们针对5个AI面试官、3个访谈主题和450名人类参与者开展了15项真实定性研究,并将其与InterviewPlayground中的模拟研究进行比较。我们发现,InterviewPlayground中AI面试官的表现与人类研究中的表现之间,在12项指标上的平均皮尔逊相关系数为0.86,且InterviewPlayground中的模拟交互重现了人类研究中AI面试官行为分析的关键发现。这些结果共同证明了InterviewPlayground在评估AI面试官性能和检查潜在失败模式方面的有效性。我们的工作贡献了一个经实证验证的AI面试官模拟环境,更广泛地为未来开发经过验证的、基于模拟的对话AI系统评估工作提供了路线图。
英文摘要
Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
CommentsPreprint. 23 pages