Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
自描述多模态交互调优:放大可利用冗余以实现鲁棒的视觉语言模型
机构 * Singapore University of Technology and Design(新加坡科技设计大学) ; DSO National Laboratories(国防部国家实验室) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 针对视觉语言模型中的幻觉和鲁棒性问题,提出自描述多模态交互调优方法,通过放大模态间冗余信息来补偿受损模态,并设计多模态交互门机制将独特交互转化为冗余交互,实验表明该方法可减少38.3%的视觉诱导错误并提升16.8%的一致性。
Comments Accepted to ICML 2026. Code: https://github.com/yurielryan/Multimodal-Interaction-Tuning