机构
*
College of Computer Science, Sichuan University(四川大学计算机学院)
;
Columbia University(哥伦比亚大学)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
Apon AI and Brain-Computer Engineering Research Institute(Apon人工智能与脑机工程研究院)
;
Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院)
;
Faculty of Applied Science and Engineering, University of Toronto(多伦多大学应用科学与工程学院)
;
Faculty of Computer Science and Information Technology, University of Malaya(马来亚大学计算机科学与信息技术学院)
;
Zhaolong Technology(智龙科技)
;
Purdue University(普渡大学)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉
Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das
机构
*
Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院)
;
TCG Crest
;
Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位)
;
Indian Statistical Institute(印度统计研究所)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Insight-A: Attribution-aware for Multimodal Misinformation Detection
Insight-A: 多模态虚假信息检测中的归因意识
Junjie Wu, Yumeng Fu, Chen Gong, Guohong Fu
机构
*
School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院)
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
Chen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu, Kejun Wu, Ruoyu Wang, Yi Wang, Soo Chin Liew
机构
*
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
The Hong Kong Polytechnic University(香港理工大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Wuhan University(武汉大学)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
University of Washington(华盛顿大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
CommentsAccepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)