Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
基于动态视觉搜索和缩放的自适应聚焦推理方法用于高效VLMs
Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li
机构
*
organization= School of Computer Science \& Technology, Beijing Institute of Technology , city= Beijing , country= China
;
organization= State Key Laboratory of General Artificial Intelligence, BIGAI , city= Beijing , country= China
;
organization= School of Intelligence Science
;
Technology, Peking University , city= Beijing , country= China
;
organization= Guangdong Laboratory of Machine Perception
;
Intelligent Computing, Shenzhen MSU--BIT University , city= Shenzhen , country= China
;
organization= Department of Automation, Tsinghua University , city= Beijing , country= China
机构
*
Key Laboratory of Opto-Electronic Information Processing, Chinese Academy of Sciences(光电信息处理重点实验室,中国科学院)
;
Shenyang Institute of Automation, Chinese Academy of Sciences(沈阳自动化研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Sun Yat-sen University(中山大学)
;
Tencent(腾讯)
;
Nankai University(南开大学)
;
MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)
机构
*
Department of Electrical Engineering
;
Computer Science University of Toledo Toledo, USA
;
Institute of Mathematical Sciences Claremont Graduate University Claremont, USA
;
Department of Bioengineering University of Toledo Toledo, USA
;
Department of Computer Science Bowling Green State University Bowling Green, USA
专题命中
领域大模型
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
机构
*
Cornell University(康奈尔大学)
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
;
University of Illinois Urbana‑Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of California, Los Angeles(加州大学洛杉矶分校)
专题命中
领域大模型
:language model(abstract);small language model(abstract);分类 cs.AI
CommentsThis update corrects a minor error in Table 1 of the originally submitted version, where a formula was inadvertently included. This does not affect the methodology or results. We also improved the formatting of existing formulas and the visual presentation of algorithm outputs. No changes were made to the scientific content
Journal refInternational Journal of Web & Semantic Technology (IJWesT) Vol.16, No.2, April 2025
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
有意识的注视:用于视觉-语言模型中幻觉抑制的自适应注意力机制
Weijue Bu, Guan Yuan, Guixian Zhang
机构
*
School of Computer Science and Technology/School of Artificial Intelligence(计算机科学与技术学院/人工智能学院)
;
China University of Mining and Technology(中国矿业大学)