发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM因系统提示词中特质不同导致同一请求安全决策不一致的问题,提出特质不变安全微调(TIST)框架及TraSN方法,提升跨特质安全稳定性且保留通用能力。
AI 中文摘要
对齐后的大语言模型(LLM)应根据用户请求内容表现出安全行为:拒绝不安全请求、遵从安全请求。但研究发现,同一请求在系统提示词中分配不同特质时,会引发差异显著的安全决策,该失效模式被称为特质诱导的安全变异。为衡量此失效,提出基于拒绝的指标:特质诱导偏差衡量数据集层面与无特质基准的偏差,特质诱导翻转率衡量同一请求在不同特质下是否获得不同安全决策。随后对特质诱导安全变化的机制开展表征层面分析,发现特质会在低维子空间内扰动模型的安全表征。为实现跨特质稳定的特质不变安全,提出特质不变安全微调(TIST),这是一种简单有效的自蒸馏框架,可将LLM的特质条件行为与其无特质行为对齐。基于分析进一步提出TIST的实例化方法——特质子空间中和(TraSN),仅在已识别的特质子空间内强制实现不变性。实验表明,TraSN可提升特质不变安全,增强有害请求的安全性,同时保留模型的通用能力。研究结果凸显特质是LLM安全与模型行为鲁棒性的重要影响因素。
英文摘要
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation. To measure this failure, we introduce refusal-based metrics: Trait-Induced Deviation measures dataset-level deviation from the no-trait baseline, while Trait-Induced Flip Rate measures whether the same request receives different safety decisions across traits. We then provide a representation-level analysis of the mechanism behind trait-induced safety shifts and find that traits perturb the model's safety representations within a low-dimensional subspace. To achieve trait-invariant safety, where safety behavior remains stable across traits, we introduce Trait-Invariant Safety Tuning (TIST), a simple yet effective self-distillation framework that aligns an LLM's trait-conditioned behavior with its no-trait behavior. Guided by our analysis, we further propose Trait-Subspace Neutralization (TraSN), an instantiation of TIST, which enforces invariance only within the identified trait subspace. Experiments show that TraSN improves trait-invariant safety and strengthens harmful-request safety while preserving general capability. Our results highlight traits as an important factor in LLM safety and robust model behavior.