arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22234cs.CLcs.CVcs.LG

透视冲突:提升视觉-语言模型中的指令层级对齐

Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models

Nicholas Sansoterra, Zishuo Zheng, Sachin Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将多模态指令层级对齐视为推理问题,用基于规则的奖励强化学习训练VLM,发现混合模态监督最佳,能提升跨模态及智能体场景下的鲁棒性并泛化到真实任务。

中文摘要 AI 辅助

指令层级(IH)对齐旨在训练语言模型在输入冲突时优先遵循更高级别的指令。虽然该研究主要在纯文本环境中进行,但视觉-语言模型(VLM)为IH引入了新的挑战:指令可能嵌入图像中、跨模态拆分、经过视觉变换,或在智能体任务中遇到。我们将多模态IH对齐视为一个推理问题,使用基于规则的奖励通过强化学习训练VLM,并比较了纯文本、纯图像和混合模态监督的效果。我们发现,纯文本IH训练能部分迁移到多模态攻击,但当模型需要跨模态解码、重建或推理指令时会失败。基于图像的训练比纯文本监督更能提升鲁棒性,而混合模态训练整体表现最佳。重要的是,这些优势超越了合成排版训练环境,泛化到真实图像和网络智能体安全任务,同时基本保留了通用多模态能力,表明轻量级、可验证的监督能显著提升VLM在对抗性、跨模态和交互式指令冲突下的鲁棒性。

英文摘要

Instruction hierarchy (IH) alignment teaches language models to prioritize higher-level instructions when inputs conflict. While studied primarily in text-only settings, vision-language models (VLMs) introduce new challenges for IH: instructions may be embedded in images, split across modalities, visually transformed, or encountered during agentic tasks. Positing multimodal IH alignment as a reasoning problem, we train VLMs using reinforcement learning with rule-based rewards, comparing text-only, image-only, and mixed-modality supervision. We find that text-only IH training partially transfers to multimodal attacks, failing when models must decode, reconstruct, or reason over instructions across modalities. Image-based training improves robustness beyond text-only supervision, while mixed-modality training performs best overall. Importantly, the benefits generalize beyond the synthetic typographic training setting to real-image and web-agent safety tasks, while largely preserving general multimodal capability, showing that lightweight, verifiable supervision can meaningfully improve VLM robustness under adversarial, cross-modal, and interactive instruction conflicts.

发表机构

  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑