HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉
Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das
机构
*
Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院)
;
TCG Crest
;
Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位)
;
Indian Statistical Institute(印度统计研究所)
机构
*
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
;
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Electronic Information School, Wuhan University(武汉大学电子信息学院)
Zhiwei Huang, Jiaqi Li, Hongbo Zhao, Xiao Ma, Ping Zhong, Xiaohu Zhou, Wei Ye, Rui Fan
机构
*
Department of Control Science & Engineering, the College of Electronic & Information Engineering, Tongji University(控制科学与工程系,电子与信息工程学院,同济大学)
;
School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)
;
Beijing Institute of Aerospace Control Devices(北京航天控制器件研究所)
;
Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)
Collaborative Attention and Consistent-Guided Fusion of MRI and PET for Alzheimer's Disease Diagnosis
Delin Ma, Menghui Zhou, Jun Qi, Yun Yang, Po Yang
机构
*
School of Software Yunnan University(软件学院 云南大学)
;
Department of Computer Science The University of Sheffield(计算机科学系 剑桥大学)
;
Department of Computing Xian JiaoTong-Liverpool University(计算系 西交利物浦大学)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi
机构
*
of Computer Science, University of Southern California, Los Angeles, CA, USA
;
Computer Engineering, University of Southern California, Los Angeles, CA, USA
Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification
Sheng Huang, Jiexuan Yan, Beiyan Liu, Bo Liu, Richang Hong
机构
*
Ministry of Education Key Laboratory of Dependable Service Computing in Cyber Physical Society(教育部可信服务计算网络社会重点实验室)
;
School of Big Data and Software Engineering(大数据与软件工程学院)
;
School of Computer Science and Information Engineering(计算机科学与信息工程学院)
VLA-Mark: A cross modal watermark for large vision-language alignment model
Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
;
University of Toronto(多伦多大学)
;
Ant Group, Alibaba(蚂蚁集团,阿里巴巴)
;
New York University Shanghai(纽约大学上海分校)
OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
Xingxin He, Aurora Rofena, Ruimin Feng, Haozhe Liao, Zhaoye Zhou, Albert Jang, Fang Liu
机构
*
Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉A.马丁诺斯生物医学成像中心)
;
Harvard Medical School(哈佛医学院)
;
Massachusetts General Hospital(麻省总医院)
;
University Campus Bio-Medico of Rome(罗马生物医学大学校园)
p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
Jun Zhang, Desen Meng, Zhengming Zhang, Zhenpeng Huang, Tao Wu, Limin Wang
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
China Mobile Research Institute(中国移动研究院)
;
Shanghai AI Lab(上海AI实验室)