arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

如何驯服多头九头蛇?面向大语言模型的自适应多类别安全引导

How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models

Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu

arXiv 2609.34514首次发表:更新:

发表机构

Southeast University; Peking University; Chongqing University(东南大学; 北京大学; 重庆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CAM-Steer框架,通过估计各危害类别风险并组合安全方向,自适应引导隐藏状态旋转,在多类别共现时提升大语言模型安全防御成功率,且推理开销极低。

AI 中文摘要

随着大语言模型(LLMs)日益普及,防止其对有害提示产生不安全响应对于其安全部署至关重要。激活引导(Activation steering)提供了一种在不更新模型参数的情况下,通过在推理过程中修改内部激活来提升LLM安全性的方法。然而,单个提示可能涉及多个危害类别,针对某一类别的安全引导可能使另一类别的有害内容未被处理。尽管自适应引导已有进展,但现有方法在单个提示中同时出现多个危害类别时,并未显式协调引导方向和强度。为解决此问题,我们提出了CAM-Steer,一种类别自适应多类别安全引导框架。具体而言,它通过将当前隐藏状态与安全和不安全原型进行比较,来估计每个危害类别相关的风险。随后,利用估计的风险将不同危害类别的安全方向组合成单一引导方向,并确定干预的强度。最后,它沿组合后的引导方向旋转隐藏状态,旋转角度由估计的风险决定,同时保持隐藏状态范数不变。在三个LLM骨干网络和七个危害类别上的实验表明,CAM-Steer在平均防御成功率上优于所评估的基线方法,包括在类别共现的情况下。进一步的分析支持其组件设计和信息丰富的风险评分,且推理开销可忽略不计。

英文摘要

As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during inference without updating model parameters. However, a single prompt can involve multiple harm categories, and steering toward safety in one category may leave harmful content from another unaddressed. Despite advances in adaptive steering, existing methods do not explicitly coordinate steering direction and strength when multiple harm categories co-occur within a single prompt. To address this problem, we propose CAM-Steer, a Category-Adaptive Multi-category Safety Steering framework. Specifically, it estimates the risk associated with each harm category by comparing the current hidden state with safe and unsafe prototypes. The estimated risks are then used to combine the safety directions for different harm categories into a single steering direction and to determine the strength of the intervention. Finally, it rotates the hidden state along the composed steering direction, with the rotation angle determined by the estimated risks, while preserving the hidden-state norm. Experiments across three LLM backbones and seven harm categories show that CAM-Steer outperforms the evaluated baselines in average defense success rate, including when categories co-occur. Further analyses support its component designs and informative risk scores, with negligible inference overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑