GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
机构 * Shandong University(山东大学)
专题命中 VLM训练与架构 :LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Shandong University(山东大学)
专题命中 VLM训练与架构 :LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * Seoul National University(首尔国立大学) ; Amazon(亚马逊)
专题命中 VLM训练与架构 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI
专题命中 VLM训练与架构 :VLM(abstract);visual language model(abstract);分类 cs.CV、cs.AI
Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community
机构 * Boston University(波士顿大学) ; Microsoft Research(微软研究院)
专题命中 VLM训练与架构 :LLaVA(abstract);multimodal large language model(abstract);分类 cs.LG
机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) ; School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) ; Nanyang Technological University(南洋理工大学) ; University of Leicester(莱斯特大学) ; School of Astronautics, Beihang University(北京航空航天大学航天学院)
专题命中 VLM训练与架构 :visual question answering(abstract);分类 cs.CV、cs.AI
机构 * Department of Diagnostic Imaging, Brown University Health(布朗大学健康中心诊断影像科) ; Warren Alpert Medical School of Brown University(布朗大学沃伦·阿尔珀特医学院) ; Department of Biomedical Engineering, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院生物医学工程系) ; Department of Radiology and Radiological Sciences, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院放射科) ; Johns Hopkins University Division of Pulmonary and Critical Care Medicine(约翰霍普金斯大学肺科与重症医学科) ; Department of Radiology, University of Colorado School of Medicine(科罗拉多大学医学院放射科)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI
机构 * International Institute of Information Technology Bangalore (IIITB), India(国际信息技术研究所(班加罗尔)) ; Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息与通信研究所(A*STAR))
专题命中 其他VLM :vision-language model(title,abstract);分类 cs.LG
Comments This is the revised and peer-reviewed version of our paper, accepted and published in the Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM 2025)
Journal ref Proc. 34th ACM International Conference on Information and Knowledge Management (CIKM), 2025
机构 * Waymo LLC(Waymo公司)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV、cs.AI