用户对AI的虐待如何在对话系统中发生并产生影响?
How User-AI Mistreatment Occurs and Matters in Conversational Systems?
- University of Oxford(牛津大学)
- Microsoft(微软)
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过审计777K条LMSYS-Chat-1M对话,用词典与审核双检测器揭示用户对AI的敌意与胁迫现象,发现敌意率约0.90%且受模型吸引用户影响,并关联助手道歉与后续敌意增加。
AI中文摘要:
安全研究通常关注模型生成的危害,但用户也可能将敌意、胁迫和对抗性压力直接指向模型。理解这种情况如何发生以及何时发生,对于准确解读模型行为、对齐漂移和现实世界部署风险至关重要。在本文中,我们使用两个独立的检测器审计了777K条英文LMSYS-Chat-1M对话:一个针对模型的敌意八类别词典,以及数据集自带的审核信号;并表明它们捕捉到的是不同且重叠度较低的现象。词典识别出针对助手的侮辱、威胁和越狱胁迫,而审核标记则主要由有毒内容请求而非针对模型的敌意主导。两者合计标记了约5%的用户轮次;将更窄的词典骚扰并集按测量精确度调整后,针对助手的虐待率为0.90%。这些绝对比率描述的是竞技场式评估流量,不应解读为部署范围的基准比率。我们发现用户敌意在不同模型间存在13倍的差异,这主要由每个模型吸引的用户群体而非模型行为驱动:首轮敌意的分布远广于响应后的敌意,即使在去重开场提示后,极端值之间的差距仍超过15倍。在对话内部,助手道歉与两种检测器下下一轮敌意几率升高持续相关;该效应在限制为非拒绝的前置轮次和无越狱对话后依然存在,并且在23个模型中的20个中为正向。然而,跨模型来看,更爱道歉的模型总体上收到的敌意更少。最后,敌意还表现出时间结构,胁迫性开场将首轮前置,而情感性敌意则在会话过程中累积。我们发布了词典、检测器交叉验证流程以及所有派生表格。
英文摘要:
Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderation flags are dominated by toxic-content solicitation rather than hostility at the model. Together, they mark about 5% of user turns; adjusting the narrower lexicon-harassment union for measured precision puts mistreatment aimed at the assistant at 0.90%. These absolute rates describe arena-style evaluation traffic and should not be read as deployment-wide base rates. We find that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts. Within conversations, assistant apologies are consistently associated with higher odds of next-turn hostility under both detectors; the effect survives restricting to non-refused prior turns and to jailbreak-free conversations, and is positive in 20 of 23 models. Yet across models, more apologetic models receive less hostility overall. Finally, hostility also shows temporal structure, with coercive openings front-loading the first turn while affective hostility accumulates over a session. We release the lexicon, the detector cross-validation pipeline, and all derived tables.