Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu
机构
*
MBZUAI
;
Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室)
;
King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
;
University of Copenhagen(哥本哈根大学)
;
Tsinghua University(清华大学)
机构
*
Institute of Trustworthy Embodied AI(可信具身人工智能研究院)
;
Fudan University(复旦大学)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
Columbia University(哥伦比亚大学)
Descriptive Image-Text Matching with Graded Contextual Similarity
Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn
专题命中
图文多模态
:image-text(title,abstract);分类 cs.CV
CommentsThis version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
York University(约克大学)
;
Mila – Quebec AI Institute(魁北克人工智能研究院)
;
École de Technologie Supérieure(魁北克高等技术学院)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Waterloo(滑铁卢大学)
;
CIFAR AI Chair(CIFAR人工智能 chair)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
University of British Columbia(不列颠哥伦比亚大学)
MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning
Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim
机构
*
School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
;
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
;
State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)
机构
*
School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computing and Information Technology, Great Bay University(大亚湾大学计算机与信息科技学院)
;
College of Computer Science, Nankai University(南开大学计算机学院)
;
School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
;
National Engineering Laboratory of Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室)
专题命中
图文多模态
:multimodal(title,abstract);分类 cs.CV
CommentsAccepted by IEEE Transactions on Instrumentation and Measurement (TIM)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian, Yoshua Bengio, Perouz Taslakian
机构
*
ServiceNow
;
Université de Montréal(蒙特利尔大学)
;
École de Technologie Supérieure(高级技术学院)
;
University of Waterloo(滑铁卢大学)
;
McGill University(麦吉尔大学)
;
York University(约克大学)
;
CIFAR AI Chair(CIFAR人工智能主席)
;
Mila
;
Law Zero
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
;
Xiamen University(厦门大学)
;
The Hong Kong University of Science and Technology(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks
Yuan Gao, Sangwook Kim, Chris McIntosh
机构
*
Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络)
;
Department of Medical Biophysics, UofT(医学生物物理学系)
;
Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络)
;
Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学)
;
Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute)
;
Department of Medical Imaging, UofT(医学影像学系)
;
Vector Institute, Toronto(向量研究所)
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
Gianmarco Spinaci Department of Classical Philology and Italian Studies, University of Bologna, Italy Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Gianmarco Spinaci 文艺复兴研究系,博洛尼亚大学,意大利 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利)
;
Lukas Klic Villa i Tatti, The Harvard University Center for Italian Renaissance Studies, Florence, Italy(Lukas Klic 塔蒂别墅,哈佛大学意大利文艺复兴研究中心,佛罗伦萨,意大利)
;
Giovanni Colavizza Department of Classical Philology and Italian Studies, University of Bologna, Italy Department of Communication, University of Copenhagen, Denmark(Giovanni Colavizza 文艺复兴研究系,博洛尼亚大学,意大利 传播系,哥本哈根大学,丹麦)
Seeing 3D Through 2D Lenses: 3D Few-Shot Class-Incremental Learning via Cross-Modal Geometric Rectification
Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
机构
*
South China University of Technology(南方科技大学)
;
State Key Laboratory of Subtropical Building Science(亚热带建筑科学国家重点实验室)
;
Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室)
;
Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室)
;
Singapore Management University(新加坡国立大学)
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
Qizao Wang, Xuelin Qian, Bin Li, Yanwei Fu, Xiangyang Xue
机构
*
School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学)
;
School of Computer Science, Shanghai Key Lab of Intelligent Information Processing, Fudan University(计算机学院,上海智能信息处理重点实验室,复旦大学)
;
Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)