发表机构
Wuhan University; The University of Tokyo; Peking University; Shanghai AI Laboratory; Shanghai Innovation Institute; Shanghai Jiao Tong University; The University of Hong Kong(武汉大学; 东京大学; 北京大学; 上海人工智能实验室; 上海创新研究院; 上海交通大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有视觉语言模型安全对齐会全局修改模型的问题,提出流式识别与门控LoRA框架,实现按需令牌级安全干预,在多基准上验证了方法的有效性。
AI 中文摘要
现有视觉语言模型的安全对齐方法通常会全局修改模型行为:一旦安全参数被训练或加载,它们会同时参与不安全和已安全的生成过程。这种持续开启的干预会不必要地干扰模型的原始推理路径,并降低通用多模态能力。我们认为安全对齐应是按需干预,而非对每个解码轨迹进行永久修改。为此,我们提出一种用于内在视觉语言模型安全的流式识别与门控LoRA框架。在自回归生成过程中,一个轻量型识别器会估计当前预令牌生成状态是否安全,其输出会更新后续解码步骤的LoRA门;否则,生成将遵循冻结主干策略。LoRA模块基于不安全前缀、转换语句和安全延续进行训练,以便在激活后学会将不安全生成重定向回安全响应。在多个安全和通用基准上开展的实验,证明了我们的方法在对齐后设置中的有效性。
英文摘要
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade general multimodal capabilities. We argue that safety alignment should be an on-demand intervention rather than a permanent modification to every decoding trajectory. To this end, we propose a streaming recognition and gated LoRA framework for intrinsic VLM safety. During autoregressive generation, a lightweight recognizer estimates whether the current pre-token generation state is safe or unsafe. Its output updates the LoRA gate for the following decoding step; otherwise, generation follows the frozen-backbone policy. The LoRA module is trained from unsafe prefixes, transition statements, and safe continuations, so that it learns to redirect unsafe generations back to safe responses after activation. Experiments across multiple safety and general-purpose benchmarks demonstrate the effectiveness of our method in post-alignment settings.
CommentsPreprint. 13 pages, 4 figures. Main paper with appendix