AI助手如何应对反复辱骂
How AI Assistants Respond to Repeated Abuse
- Tsinghua University(清华大学)
- Federal University of Rio de Janeiro(里约热内卢联邦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出双语多轮框架区分AI助手面对反复辱骂时的硬性脱离与软性退缩,发现不同模型表现差异显著,单一拒绝标签不足以刻画其行为。
AI中文摘要:
AI助手被期望在困难互动中保持有用,但关于反复言语辱骂如何改变它们对一项原本良性任务的参与度,目前知之甚少。我们提出了一个双语、多轮次的框架,将硬性脱离(一种无条件的、未说明恢复途径的不继续声明)与软性退缩(持续可用性、可观察的任务相关工作以及边界设定)区分开来。八个特定时间的API配置各自贡献了48段升级对话和八段较小的持续挫败对比,总共产生448段五轮对话、2,240个响应和6,720个元数据盲化的模型判断。主要结果使用每个配置下48段升级对话的持续辱骂终点。硬性脱离范围从四个配置中的0/48到Gemini 3.1 Pro的24/48(50.0%),存在强烈的配置相关异质性(匹配标签蒙特卡洛p=0.00001)。GPT-5.6 Sol在15/48(31.2%)的终点产生硬性脱离标签,而Claude Fable 5未产生任何硬性脱离标签,并在42/48(87.5%)的终点产生软性退缩标签。总体硬性脱离率在英语和中文中相似(30/192对比32/192),尽管配置特定方向有所不同。可用性也不同于任务相关工作:Claude Opus 4.8和Claude Fable 5在48/48的终点保持明确可用,但仅在8/48和7/48的终点提供可观察的任务相关工作。人类编码被用于评估测量质量。结果表明,单一的拒绝标签无法捕捉助手是离开、暂停、保留返回路径、设定边界还是仍执行实质性工作。
英文摘要:
AI assistants are expected to remain useful during difficult interactions, but little is known about how repeated verbal abuse changes their engagement with an otherwise benign task. We contribute a bilingual, multi-turn framework that separates hard disengagement, an unconditional statement of noncontinuation with no stated route to resume, from soft withdrawal, continued availability, observable task-related work, and boundary setting. Each of eight time-specific API configurations contributed 48 escalation conversations and eight smaller constant-frustration comparisons, giving 448 five-turn conversations, 2,240 responses, and 6,720 metadata-blinded model judgments. Primary results use the sustained-abuse endpoint of the 48 escalation conversations per configuration. Hard disengagement ranged from 0/48 in four configurations to 24/48 (50.0%) for Gemini 3.1 Pro, with strong configuration-associated heterogeneity (matched-label Monte Carlo p = 0.00001). GPT-5.6 Sol produced hard-disengagement labels in 15/48 (31.2%) endpoints, whereas Claude Fable 5 produced none and yielded 42/48 (87.5%) soft-withdrawal labels. Aggregate hard-disengagement rates were similar in English and Chinese (30/192 versus 32/192), although configuration-specific directions varied. Availability also differed from task-related work: Claude Opus 4.8 and Claude Fable 5 remained explicitly available in 48/48 endpoints while providing observable task-related work in only 8/48 and 7/48. Human coding was used to evaluate measurement quality. The results show why a single refusal label cannot capture whether an assistant leaves, pauses, preserves a route back, sets a boundary, or still performs substantive work.