发表机构
Southern University of Science and Technology; Sun Yat-sen University; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); University of Electronic Science and Technology of China; Hangzhou Dianzi University; Shenzhen MSU-BIT University; GigaAI; The Chinese University of Hong Kong (Shenzhen)(南方科技大学; 中山大学; 广东省人工智能与数字经济实验室(深圳); 电子科技大学; 杭州电子科技大学; 深圳北理莫斯科大学; 极佳科技; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出意图特权OPSD方法,利用意图监督训练VLM识别隐性风险,以少量数据实现安全与有用性的高效对齐,显著提升联合成功率。
AI 中文摘要
视觉语言模型(VLMs)仍然容易受到跨模态隐性风险的影响:单独看似良性的视觉和文本输入,共同作用时可能引发不安全响应。现有安全方法通常需要大规模偏好数据集、昂贵的多次展开训练,或在推理时增加额外防护措施。它们也可能因直接拒绝本可安全回答的请求而牺牲有用性。本文提出意图特权在线策略自蒸馏(Intent-Privilege On-Policy Self-Distillation, OPSD),利用基于证据的意图作为训练期间的监督特权,帮助VLM识别隐性风险并提供安全、有用的响应,而非一概拒绝。OPSD通过每个提示仅一次展开,将教师基于意图的响应偏好蒸馏给学生;学生随后无需意图标注或额外安全模块即可响应。仅使用1,447个安全特定示例——比标准偏好数据集少95%——OPSD相对于多次展开的GRPO式训练将训练时间减少5倍,平均推理长度减少7%。在所有五个评估组中,OPSD取得了联合安全-有用性成功率的最高比例,该指标衡量既安全又有用的响应占比。值得注意的是,在合并的SIUO+HoliSafe上,该成功率从43.9%提升至53.5%。这些结果表明,训练时的意图监督能在大幅降低数据、训练和推理成本的同时,提升安全性和有用性。
英文摘要
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.