TRACE:面向多轮对抗对话评估的轨迹感知推理
TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation
查看机构详情
- Texas A&M University(德克萨斯农工大学)
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对LLM多轮越狱攻击,提出具备轨迹感知推理的Trace防御方法,训练Llama-3.1-8B-Instruct实现安全与实用平衡,在多轮攻击基准中显著降低攻击成功率并提升合规率。
中文摘要 AI 辅助
多轮越狱攻击已成为大型语言模型(LLM)面临的关键安全威胁,攻击者会将有害目标拆解为一系列看似良性的对话轮次,以此绕过安全护栏。现有防御措施缺乏识别演变操纵模式的推理能力,往往因过度拒绝与敏感话题相关的良性请求而牺牲实用性以换取安全性。我们提出了Trace,一种具备轨迹感知结构化推理能力的多轮防御方法。在生成每一轮响应前,该模型会从对话轨迹中识别操纵线索,评估用户意图的良性与对抗性两种解读,分配越狱评分,并确定执行以下动作之一:允许、谨慎处理或弃权(不执行)。我们从五个攻击框架中整理了4000个多轮对抗对话,搭配2400个良性对话及600个敏感但良性的对话。我们采用SFT与GRPO训练Llama-3.1-8B-Instruct模型,使用多组件奖励函数,该函数同时优化对良性提示的实用性及抵御越狱尝试的鲁棒性。在七个多轮攻击基准测试中,Trace的平均攻击成功率(ASR)为14.5%,而最强基线的平均ASR为31.4%,未防御目标的平均ASR为74.9%,同时显著提高了每次成功越狱所需的攻击者工作量。Trace还平衡了可用性与安全性,在过度拒绝基准测试中达到93.3%的平均合规率。
英文摘要
Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model identifies manipulation cues from the trajectory, evaluates both the benign and adversarial interpretations of user intent, assigns a jailbreak score, and commits to an action: Allow, Caution, or Decline. We curate 4k multi-turn adversarial conversations from five attack frameworks, pair them with 2.4k benign dialogs, and 600 sensitive-but-benign conversations. We train Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts. Across seven multi-turn attack benchmarks, Trace attains an average attack success rate (ASR) of 14.5% against 31.4% for the strongest baseline and 74.9% for the undefended target, while significantly raising the attacker effort required per successful jailbreak. Trace also balances usability and safety, achieving a 93.3% average compliance on over-refusal benchmarks.