arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18628cs.CVcs.CL

当安全覆盖视觉时:探索视觉语言模型中视觉影响与安全对齐之间的动态关系

When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models

Mehak Gupta, Tanmoy Chakraborty

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究对齐视觉语言模型中安全诱导弃权的内部解码动态,发现安全对齐未抑制感知接地,抑制拒绝相关表征可恢复接地回答,揭示了其一种新的失败模式。

中文摘要 AI 辅助

对齐的视觉语言模型(VLMs)旨在平衡基于视觉的推理与安全生成行为。然而,我们观察到一个显著现象:在受安全约束的指令下,模型会频繁对那些在默认指令下仍可正确回答的问题作出弃权(不执行)的回应,即便输入的图像与问题完全相同。这引发了一个根本性问题:安全对齐是否会抑制感知接地本身,还是视觉证据在内部仍可用,只是生成被转向了弃权(不执行)?在本研究中,我们调查了对齐VLMs中安全诱导弃权(不执行)背后的内部解码动态。在多种架构和多模态基准上,我们表明弃权(不执行)的生成在整个解码过程中始终受到视觉证据的影响,这表明尽管存在拒绝行为,感知接地在很大程度上得以保留。我们进一步证明,尽管不同架构中拒绝的表征组织存在显著差异,但受安全约束的指令始终会将后期隐藏状态动态转向拒绝导向的解码。最后,通过有针对性的激活层面干预,我们表明抑制与拒绝相关的表征可在不重新训练或修改视觉输入的情况下,可靠地恢复跨模型的接地回答行为。这些发现共同揭示了对齐VLMs中一种此前未被充分探索的失败模式:即使感知证据在内部仍被保留,安全对齐也可覆盖基于视觉的接地表达。

英文摘要

Aligned vision-language models (VLMs) are designed to balance grounded visual reasoning with safe generation behavior. However, we observe a striking phenomenon: under safety-constrained instruction, models frequently abstain from answering questions that remain correctly answerable under default instruction despite receiving identical image-question inputs. This raises a fundamental question: does safety alignment suppress perceptual grounding itself, or does visual evidence remain internally available while generation is redirected toward abstention? In this work, we investigate the internal decoding dynamics underlying safety-induced abstention in aligned VLMs. Across multiple architectures and multimodal benchmarks, we show that abstained generations remain consistently influenced by visual evidence throughout decoding, indicating that perceptual grounding is largely preserved despite refusal behavior. We further demonstrate that, although the representational organization of refusal differs substantially across architectures, safety-constrained instruction consistently alters late-stage hidden-state dynamics toward refusal-oriented decoding. Finally, through targeted activation-level interventions, we show that suppressing refusal-related representations reliably restores grounded answering behavior across models without retraining or modifying visual inputs. Together, these findings reveal a previously underexplored failure mode in aligned VLMs: safety alignment can override grounded visual expression even when perceptual evidence remains internally preserved.

↑