LongGuard:针对安全护栏长上下文失效的机制分析与无训练缓解方案
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
浏览论文内容
中文总结 AI 辅助
本研究针对LLM安全护栏在长上下文下的失效问题,提出LongGuard框架,经机制分析后开发无训练缓解方案,在多基准测试中实现显著性能提升。
中文摘要 AI 辅助
安全护栏是防范大语言模型(LLM)有害输入输出的最后一道防线,但其训练与评估几乎仅基于短文本。本文提出LongGuard框架,用于评估、机制分析并缓解长上下文下的护栏失效问题。我们将该任务建模为长度范围0.25k至32k的“安全针在草垛中”(SafetyNIAH)任务;对15种主流护栏的实验显示,不安全内容召回率平均单调下降超50%,且通过“良性填充vs针重复”的配对设计,将失效归因于不安全针的比例稀释,而非绝对长度。对6种护栏的三层注意力- logit-行为分析定位了机制:不安全针的注意力质量被稀释,不安全内容相对于安全内容的logit边际同步压缩,检测决策随之崩溃,且在排除长度影响后,该注意力→logit→行为链仍保持一致。我们进一步分离出一组相对于基础模型具有部分特异性的护栏专用检索头。基于上述分析,我们提出两种无训练缓解方案——分块检测(Chunked Detection,CD)与注意力头锐化(Attention-Head Sharpening,AHS),以及一种部署协议——上下文感知超参数路由(Context-Aware Hyperparameter Routing,CAHR),该协议可根据上下文长度和审计侧选择配置。在涵盖合成数据、长上下文攻击及推理模型输出的5个基准测试中,CAHR-CD与CAHR-AHS分别使6种护栏的平均性能提升22%与13%。代码与数据已在线公开。
英文摘要
Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanistically analyzes, and mitigates long-context guardrail failure. We formulate the task as Safety Needle-in-a-Haystack (SafetyNIAH) over a 0.25k-32k length grid; across 15 mainstream guardrails, unsafe recall drops monotonically by more than 50% on average, and a paired Benign-Fill vs. Needle-Repeat design attributes the failure to proportional dilution of the unsafe needle rather than to absolute length. A three-layer attention-logit-behavior analysis on six guardrails locates the mechanism: attention mass on the unsafe needle is diluted, the unsafe-over-safe logit margin is compressed in lockstep, and the detection decision collapses accordingly, with this attention->logit->behavior chain remaining consistent after partialling out length. We further isolate a sparse set of guard-specialized retrieval heads that exhibit partial specificity relative to their base models. Building on the analysis, we propose two training-free mitigations - Chunked Detection (CD) and Attention-Head Sharpening (AHS) - and a deployment protocol, Context-Aware Hyperparameter Routing (CAHR), that selects configurations by context length and audit side. Across five benchmarks spanning synthetic data, long-context attacks, and reasoning-model outputs, CAHR-CD and CAHR-AHS improve the six-guardrail average by 22% and 13%, respectively. Code and data are available online.
发表机构
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
- School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。