V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models
V-SEAM:视觉语义编辑与注意力调节用于视觉-语言模型的因果可解释性
机构 * Tongji University(同济大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
AI总结 本文提出V-SEAM框架,结合视觉语义编辑和注意力调节,提升视觉-语言模型的因果可解释性,通过多级语义分析增强模型性能。
Comments EMNLP 2025 Main