Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
揭示视觉-语言模型的脆弱性:通过纹理约束扰动和跨模态优化的多模态对抗协同
Xiang Fang, Wanlong Fang, Changshuo Wang
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
Comments12 pages, 3 figures, accepted at ICMHI 2026, 10th International Conference on Medical and Health Informatics, Kyoto, Japan. To appear in ACM Conference Proceedings
机构
*
School of Intelligence Science and Technology(智能科学与技术学院)
;
State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室)
;
Nanjing University(南京大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.AI、cs.LG
CommentsThis paper has been accepted by the International Journal of Computer Vision (IJCV), 2026. The first two authors contributed equally to this work. 28 pages
机构
*
McGill University(麦吉尔大学)
;
Mila - Quebec AI Institute(魁北克人工智能研究所)
;
University of Cambridge(剑桥大学)
;
MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学)
;
University of Toronto(多伦多大学)
;
Salesforce