发表机构
King Abdullah University of Science and Technology; University of Macau; RIKEN Center for Advanced Intelligence Project; University of Electronic Science and Technology of China(阿卜杜拉国王科技大学; 澳门大学; 理化学研究所先进智能研究中心; 电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TRACE通过令牌级回报归因与对比擦除,将单轮安全抑制扩展到多轮轨迹,在35个模型-攻击组合中实现最低攻击成功率,同时保持模型效用。
AI 中文摘要
经过安全对齐的大型语言模型(LLMs)通常会拒绝有害请求,但一旦同一目标分散在多个轮次中,模型可能会遵从。偏好目标对单个提示的完整响应进行评分,因此仅靠其训练损失无法控制未见历史轨迹上的风险。我们的分析给出了充分条件,在这些条件下,在监督式单轮上下文中的抑制能够产生对多轮轨迹风险的上界。该上界考虑了覆盖度、迁移差距和泄漏,并刻画了相对于在训练策略上下文上评估的基线策略风险预算的收缩。TRACE(轨迹回报归因与对比擦除)将此原则转化为一个令牌级目标。在安全响应上,每个令牌按拒绝归因优势的折现回报进行加权。该优势将冻结的参考模型与其拒绝消融副本进行比较,使得较早的响应令牌能够从较晚的拒绝相关证据中获得信用。在拒绝响应的高差距位置,TRACE将观察到的令牌与策略选择的替代项结合在擦除目标中。梯度范数惩罚取代了保留集。在五个开放权重模型和七种多轮攻击中,TRACE在所有35个模型与攻击组合中实现了最低的攻击成功率(ASR),而在MMLU和HellaSwag上评估的模型效用最多下降1.23个百分点。源代码可在补充材料中找到。
英文摘要
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1\.23 points. Source code can be found in the supplemental material.