发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过VR模拟实验发现,警察对黑人男性虚拟角色的说话恭敬程度更低,还探讨了LLM在ATE估计中的应用,推荐混合效应模型结合iptw方法,认为LLM相关技术需进一步完善。
AI 中文摘要
在美国警察与公众互动的暴力背景下,本研究通过虚拟现实(VR)模拟,探究警察对被描绘为黑人成年男性的虚拟角色的说话方式。我们从因果推断视角评估此类角色的影响,其中将黑人男性角色分配给警察及模拟场景作为处理变量。边际平均处理效应(ATE)衡量该角色对警察每轮对话表述恭敬程度的社会影响。令人警醒的是,多数警察对黑人男性角色的说话恭敬程度更低,白人、混血及多种族女性警察除外,尤其在已知VR角色为嫌疑人的场景中。在典型VR场景的完整对话中,这些边际ATE可导致语气恭敬程度出现显著变化(0-10分制下相差2至数个点),超出感知黑人男性角色带来的初始影响。更令人不安的是,这可能导致对话破裂,进而可能引发针对公众和警察的暴力或危险情况。我们还探究了大语言模型(LLM)在ATE估计中的能力。通过方法比较分析,包括针对合成数据的模型验证,我们对LLM辅助的ATE估计方法提供了独特的科学见解。因此,对于含文本的多级数据的ATE估计,我们推荐混合效应模型结合逆倾向加权(iptw)方法,该方法利用LLM进行文本特征创建。尽管我们也测试了用于ATE估计的微调预测模型的LLM,但结论是它们仍需进一步开发和完善。
英文摘要
Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through a causal in- ference lens, where the assignment of the Black man character to a police officer and simulation is the treatment variable. Our (marginal) average treatment effect AT E measures the social impact of the character on the deference of officer statements with each turn of the conversation. Soberingly, we find that most officers speak less deferentially to Black man characters, except for White, biracial, and multiracial female officers, es- pecially in settings where the VR character was known to be a suspect. Across a full conversation of a typical VR scene, these marginal AT Es can result in notable changes in def- erence of tone (two to several points difference on a scale of 0-10), above and beyond that due to the initial effect of per- ceiving a Black male character. Even more disconcerting is that this can contribute to conversation breakdowns that po- tentially result in violence or danger to both the public and the police. We also explored the capabilities of large language models (LLMs) for ATE estimation. From our methods com- parison analysis, including model validation against synthetic data, we provide unique scientific insights on LLM-assisted methodologies for ATE estimation. As such, for ATE esti- mation with multilevel data with text, we recommend mixed effects models with the inverse propensity treatment weighted (iptw) approach, which utilized an LLM for text feature cre- ation. While we also tested LLMs for finetuning prediction models ultimately for ATE estimation, we conclude they are an area for further development and refinement.