通过后验重加权理解多模态上下文越狱
Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting
浏览论文内容
中文总结 AI 辅助
本文提出后验重加权框架,将多模态大模型上下文越狱建模为证据累积过程,导出缩放定律,并据此设计后验感知推理时防御,在固定干预预算下显著提升鲁棒性-效用权衡。
中文摘要 AI 辅助
上下文学习(ICL)越狱揭示了多模态大语言模型(MLLMs)的一个关键漏洞:提示中的有害示例可以在不修改模型参数的情况下诱导不安全输出。尽管有大量经验证据,现有工作缺乏对这类越狱为何能可靠成功或其有效性如何随上下文组成扩展的原则性理解。我们提出一个后验重加权框架,将安全对齐的MLLM建模为隐式地在竞争行为模式上运行,并将上下文示例解释为推理时证据,动态地在安全与有害行为之间转移模型的后验偏好。这一视角将越狱形式化为证据累积过程,产生了关于示例数量、有害比例、对抗强度和语义多样性的预测性缩放定律。在该框架指导下,我们引入一种后验感知的推理时防御,基于估计风险自适应注入良性反证据,有效抑制有害后验漂移,同时保持模型效用。与现有上下文防御相比,我们的方法在固定干预预算下实现了显著改进的鲁棒性-效用权衡。综上,我们的结果确立了后验重加权作为理解和缓解MLLMs中ICL越狱的统一且预测性框架。
英文摘要
In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite extensive empirical evidence, existing work lacks a principled understanding of why such jailbreaks reliably succeed or how their effectiveness scales with context composition. We propose a posterior reweighting framework that models a safety-aligned MLLM as implicitly operating over competing behavioral modes, and interprets in-context demonstrations as inference-time evidence that dynamically shifts the model's posterior preference between safe and harmful behaviors. This view formalizes jailbreak as a process of evidence accumulation, yielding predictive scaling laws with respect to demonstration count, harmful ratio, adversarial strength, and semantic diversity. Guided by this framework, we introduce a posterior-aware inference-time defense that adaptively injects benign counter-evidence based on estimated risk, effectively suppressing harmful posterior drift while preserving model utility. Compared to existing in-context defenses, our method achieves a significantly improved robustness-utility trade-off under a fixed intervention budget. Together, our results establish posterior reweighting as a unifying and predictive framework for understanding and mitigating ICL jailbreak in MLLMs.