MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
MMT-ARD: 多模态多教师对抗蒸馏用于鲁棒视觉-语言模型
Yuqi Li, Junhao Dong, Chuanguang Yang, Shiping Wen, Piotr Koniusz, Tingwen Huang, Yingli Tian, Yew-Soon Ong
机构
*
The City University of New York, CUNY(纽约城市大学)
;
Nanyang Technological University(南洋理工大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Technology Sydney(悉尼技术大学)
;
Data61, CSIRO(CSIRO数据61研究所)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu
机构
*
MBZUAI
;
Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室)
;
King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
;
University of Copenhagen(哥本哈根大学)
;
Tsinghua University(清华大学)
机构
*
Institute of Trustworthy Embodied AI(可信具身人工智能研究院)
;
Fudan University(复旦大学)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
Columbia University(哥伦比亚大学)
Descriptive Image-Text Matching with Graded Contextual Similarity
Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn
专题命中
图文多模态
:image-text(title,abstract);分类 cs.CV
CommentsThis version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
York University(约克大学)
;
Mila – Quebec AI Institute(魁北克人工智能研究院)
;
École de Technologie Supérieure(魁北克高等技术学院)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Waterloo(滑铁卢大学)
;
CIFAR AI Chair(CIFAR人工智能 chair)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
University of British Columbia(不列颠哥伦比亚大学)
MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning
Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim
机构
*
School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
;
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
;
State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)
机构
*
School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院)
;
College of Computer Science, Nankai University(南开大学计算机学院)
;
School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
;
National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CV
CommentsAccepted by IEEE Transactions on Instrumentation and Measurement (TIM)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian
机构
*
ServiceNow
;
Université de Montréal(蒙特利尔大学)
;
École de Technologie Supérieure(高级技术学院)
;
University of Waterloo(滑铁卢大学)
;
McGill University(麦吉尔大学)
;
York University(约克大学)
;
CIFAR AI Chair(CIFAR人工智能主席)
;
Mila
;
Law Zero
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
;
Xiamen University(厦门大学)
;
The Hong Kong University of Science and Technology(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks
Yuan Gao, Sangwook Kim, Chris McIntosh
机构
*
Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络)
;
Department of Medical Biophysics, UofT(医学生物物理学系)
;
Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络)
;
Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学)
;
Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute)
;
Department of Medical Imaging, UofT(医学影像学系)
;
Vector Institute, Toronto(向量研究所)
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利)
;
Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利)
;
Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)