XAI-CLIP: ROI-Guided Perturbation Framework for Explainable Medical Image Segmentation in Multimodal Vision-Language Models
XAI-CLIP: 通过区域感兴趣引导扰动框架实现多模态视觉-语言模型中可解释的医学图像分割
Thuraya Alzubaidi, Sana Ammar, Maryam Alsharqi, Islem Rekik, Muzammil Behzad
机构
*
King Fahd University of Petroleum and Minerals(国王法赫德石油和矿物大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Imperial College London(伦敦帝国学院)
;
KFUPM-SDAIA Joint Research Centre for Artificial Intelligence(KFUPM-SDAIA联合人工智能研究中心)
CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
CoTZero:通过分层合成CoT实现无标注的人类级视觉推理
Chengyi Du, Yazhe Niu, Dazhong Shen, Luxin Xu
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong MMLab(香港中文大学 MMLab)
;
The College of Computer Science and Technology(计算机科学与技术学院)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
通过LLM赋能的视觉指令微调提升几何推理能力
Zhihao Li, Yao Du, Yang Liu, Yan Zhang, Yufang Liu, Mengdi Zhang, Xunliang Cai, Charles Ling, Boyu Wang
机构
*
Department of Computer Science, Western University(计算机科学系,西部大学)
;
Meituan Inc.(美团公司)
;
Department of Automation, Tsinghua University(自动化系,清华大学)
;
School of Computer Science and Technology, East China Normal University(计算机科学与技术学院,东华大学)