arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReFrame:多模态大语言模型中基于证据的测试时安全对齐

ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models

Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai, Dawei Feng, Huaimin Wang

arXiv 2608.21100首次发表:更新:

发表机构

College of Computer Science and Technology, National University of Defense Technology; State Key Laboratory of Complex & Critical Software Environment; School of Computer Science, Wuhan University(国防科技大学计算机学院; 复杂与关键软件环境国家重点实验室; 武汉大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有多模态安全对齐方法不适用于闭源模型的问题,提出无需训练的ReFrame框架,通过两个智能体协作实现测试时安全对齐,在提升安全性能的同时保留多模态效用。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models, MLLMs)将模型能力扩展至文本之外,但也使安全对齐愈发具有挑战性。多模态安全对齐方法需应对跨模态越狱、安全意识缺失及过度敏感拒绝等问题。然而现有方法常依赖再训练或内部状态检查,限制了其在已部署闭源MLLMs中的适用性,从而催生了测试时安全对齐的需求。我们分析该场景并确定两个关键障碍:效用主导性与推理惯性,二者会导致模型忽视潜在风险或遵循恶意推理轨迹。基于这些见解,我们提出ReFrame,这是一个无需训练的多模态输入重构框架,其中两个智能体共享一个轻量级本地部署的MLLM:证据生成智能体构建互补的风险与效用证据,重写与路由智能体将其转换为安全代理提示及图像路由决策,随后调用下游MLLM,全程无需修改下游模型或访问其内部信息。在多个MLLMs及基准上开展的实验表明,ReFrame在提升越狱防御能力、安全意识及降低过度敏感性的同时,保留了多模态效用。

英文摘要

While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods must address cross-modal jailbreaks, safety-awareness failures, and over-sensitive refusals. However, existing methods often rely on retraining or internal-state inspection, limiting their applicability to deployed closed-source MLLMs and motivating test-time safety alignment. We analyze this setting and identify two key obstacles, utility dominance and reasoning inertia, which cause models to overlook latent risks or follow malicious reasoning trajectories. Guided by these insights, we propose ReFrame, a training-free multimodal input reframing framework where two agents share a lightweight locally deployed MLLM: the evidence-generation agent constructs complementary risk and utility evidence, and the rewrite-and-routing agent converts it into a safe proxy prompt and image-routing decision before calling the downstream MLLM, without modifying it or accessing its internal information. Experiments across multiple MLLMs and benchmarks show that ReFrame improves jailbreak defense, safety awareness, and oversensitivity reduction while preserving multimodal utility.

CommentsAccepted to EMNLP 2026. 22 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑