ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
ReFrame:多模态大语言模型中基于证据的测试时安全对齐
机构 * College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院) ; State Key Laboratory of Complex & Critical Software Environment(复杂与关键软件环境国家重点实验室) ; School of Computer Science, Wuhan University(武汉大学计算机学院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.AI
AI总结 该研究针对现有多模态安全对齐方法不适用于闭源模型的问题,提出无需训练的ReFrame框架,通过两个智能体协作实现测试时安全对齐,在提升安全性能的同时保留多模态效用。
Comments Accepted to EMNLP 2026. 22 pages, 7 figures, 5 tables