多轮人机对话中的善意偏见
Benevolent Bias in Multi-Turn Human-Agent Dialogue
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对人机对话中隐藏在积极语气后的善意偏见,构建了标注语料库并测试两类检测器,发现现有安全检测器难识别善意偏见,LLM评判器易误判,需关注对待方式差异以实现公平监控。
AI中文摘要:
人机交互中的偏见不仅可通过敌对语言体现,还可表现为善意偏见——即不平等对待隐藏在温暖积极的语气背后。为实现此类偏见的可检测性,我们从语气和对待方式两个维度对善意偏见进行操作化定义,划分出三类:中性支持、显性偏见和善意偏见。基于这些定义,我们构建了BENEVDIAL,这是一个包含362880段多轮支持对话的类别平衡语料库,涵盖用户与智能体的人口统计特征、角色及生成者,以支持受控评估。随后我们在该语料库上测试了两类检测器:现成的安全检测器和经提示的大型语言模型(LLM)评判器。结果显示存在检测差距:现成检测器能可靠识别显性偏见,但基本遗漏善意偏见;LLM评判器在更明确的检测标准下能捕捉到更多善意偏见,但越来越多地将中性支持误分类为善意偏见,且人口统计背景会放大误报。这些发现表明,人机对话的公平监控必须超越表面线索,关注智能体的对待方式是否存在差异。
英文摘要:
Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our findings reveal a notable detection gap: off-the-shelf detectors reliably flag overt bias yet largely fail to identify benevolent bias. LLM judges improve sensitivity when guided by explicit detection criteria, but this comes at the cost of increased misclassification of neutral supportive statements as benevolent bias, a tendency that is further exacerbated by the presence of demographic context. These findings suggest that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate.