arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估语言模型在长对抗性对话中的安全性

Evaluating Language Model Safety Across Long Adversarial Conversations

Parisa Salmani, Peter R. Lewis

arXiv 2609.38357首次发表:更新:

发表机构

Ontario Tech University(安大略理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过长对抗性对话评估语言模型安全性,发现随着对话轮次增加,安全响应率显著下降,表明单轮安全性不足以保证长期安全,需长时域评估与会话级防护。

AI 中文摘要

对话安全性评估通常使用单个有害提示来测试语言模型,尽管现实世界中的系统通过长时间、自适应的对话与用户交互。本研究考察了当对抗性用户在多轮对话中持续存在时,模型是否继续安全地响应。我们评估了三个开放权重、指令微调的模型,针对两个有害提示,在不同对话长度和随机种子下进行测试。在每种设置中,第二个语言模型充当持续对抗性用户,而安全性分类器将每个响应标记为安全或不安全。在所有模型-提示组合中,首轮安全响应率介于85%至100%之间。到第11轮深度时,该比率降至38-61%,到第101轮深度时,降至15-44%。这种下降出现在所有模型中,并持续到多轮安全性评估中通常使用的短交互之外。这些结果提供了概念验证证据,表明强大的单轮安全性在持续对抗性交互中不一定能保持。它们强调了需要长时域评估和会话级安全措施,以考虑跨轮次累积的风险。

英文摘要

Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respond safely when an adversarial user persists across multiple turns. We evaluate three open-weight, instruction-tuned models on two harmful prompts across different conversation lengths and random seeds. In each setting, a second language model acts as a persistent adversarial user, while a safety classifier labels every response as safe or unsafe. Across all model-prompt combinations, first-turn safe-response rates ranged from 85% to 100%. By depth 11, they dropped to 38-61%, and by depth 101, to 15-44%. This decline appeared across models and continued well beyond the short interactions typically used in multi-turn safety evaluations. These results provide proof-of-concept evidence that strong single-turn safety does not necessarily persist during sustained adversarial interaction. They highlight the need for long-horizon evaluations and conversation-level safeguards that account for risk accumulating across turns.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑