arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10513cs.CVcs.AI

SafeCap:通过图像字幕强化学习提升大型视觉语言模型(LVLM)的安全性

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Caoyuan Ma, Wenpu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Yinqiang Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

SafeCap是一种通过图像字幕强化学习提升LVLM安全性的框架,在多模态安全基准上安全得分提升3.7-19.0个百分点,且视觉效用相当或更优,性能优于安全SFT等方法。

中文摘要 AI 辅助

大型视觉语言模型(LVLM)仍易受越狱攻击,此类攻击利用视觉输入绕过其语言主干继承的安全对齐。我们提出SafeCap,这是一种通过学习自字幕实现LVLM安全对齐的强化学习框架。SafeCap训练一个策略模型,先生成与安全相关的图像字幕,再生成最终答案;该字幕会通过冻结大语言模型(LLM)是否能达成安全对齐决策来进一步优化。这种字幕介导的目标鼓励策略模型暴露与安全响应生成相关的视觉线索,而非仅依赖直接拒绝监督。在5个多模态安全基准和6个视觉效用基准上,SafeCap在其指定的DirectCap协议下大幅提升了整体安全性能,在4种模型设置下安全平均得分提升了3.7至19.0个百分点,同时保持相当或更优的视觉效用。在匹配主干和数据的对照实验中,SafeCap的表现优于安全SFT、DPO和SafeGRPO,证明了字幕介导的强化学习在多模态安全对齐中的有效性。

英文摘要

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.

补充信息

↑