Suppressing VLM Hallucinations with Spectral Representation Filtering
机构 * Tel Aviv University(特拉维夫大学)
专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);LLaVA(abstract);grounding(abstract)
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Tel Aviv University(特拉维夫大学)
专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);LLaVA(abstract);grounding(abstract)
机构 * Tsinghua University(清华大学) ; Beijing Jiaotong University(北京交通大学) ; Inspur Yunzhou Industrial Internet Co., Ltd(Inspur云洲工业互联网有限公司)
专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV
机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) ; Independent Researcher(独立研究者) ; Duke University(杜克大学) ; University of Michigan Ann Arbor(密歇根大学安娜堡分校) ; Georgia Institute of Technology(佐治亚理工学院) ; University of California, Irvine(加州大学 Irvine 分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments 6pages,3 figures.Uunder review of International Conference on Artificial Intelligence, Computer, Data Sciences and Applications
机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) ; University of Toronto(多伦多大学) ; University of Notre Dame(诺特大学) ; Stony Brook University(石溪大学)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * College of Information Science(信息科学学院) ; Electronic Engineering, Zhejiang University, Hangzhou, 310027, Zhejiang, China(电子工程系,浙江大学,杭州,310027,浙江,中国)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG
Comments Submitting for Neurocomputing
机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Department of Radiology, Renmin Hospital of Wuhan University(武汉大学仁医院放射科)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments 40 pages
机构 * SJTU(上海交通大学) ; AgiBot ; Shanghai AI Lab(上海人工智能实验室) ; CUHK MMLab(香港大学多模态实验室) ; LV-NUS Lab(南洋理工大学-立陶宛实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG
Comments Accepted by NeurIPS 2025. Website: https://sites.google.com/view/enerverse
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments 10 pages, 3 figures, RAGE-KG 2025
机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) ; Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments 11 pages
机构 * University of Science and Technology of China(中国科学技术大学) ; The University of Hong Kong(香港大学) ; King’s College London(伦敦国王学院) ; Peking University(北京大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
专题命中 文档图表理解 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Korea University(韩国大学)
专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments AAAI 2026 (Main Technical Track)
机构 * CVIT, Kohli Centre for Intelligent Systems, IIIT Hyderabad(IIIT海得拉尔计算机视觉研究所、Kohli智能系统中心) ; Adobe Research, Bengaluru(Adobe研究)
专题命中 文档图表理解 :vision-language model(abstract);分类 cs.CV
Comments 25 pages, 21 figures
机构 * Samsung SDS(三星SDS)
专题命中 GUI与屏幕智能体 :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments 26 pages, 7 figures. Code available at https://github.com/samsungsds-research-papers/mega-gui
专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV
专题命中 GUI与屏幕智能体 :MLLM(title,abstract);multimodal large language model(abstract)
机构 * Waseda University(早稻田大学) ; CyberAgent, Inc.(CyberAgent公司) ; AI Shift, Inc.(AI Shift公司) ; Nara Institute of Science and Technology(奈良研究所)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
Comments AAAI 2026
机构 * School of Electrical Engineering, KAIST(韩国科学技术院电子工程学院)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV
Comments 9 pages, 9 figures
机构 * MAIS, Institute of Automation, Chinese Academy of Sciences, China(自动化研究所,中国科学院,中国) ; School of Artificial Intelligence, University of Chinese Academy of Sciences, China(中国科学院大学人工智能学院,中国) ; Alibaba Group(阿里巴巴集团)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.AI
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.AI
Comments 18 pages, 11 figures
机构 * Xiaomi EV(小米电动车) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallagher人工智能学院) ; School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)
专题命中 GUI与屏幕智能体 :vision-language model(abstract)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Accepted by IEEE Transactions on Multimedia
机构 * Nimblemind USA(Nimblemind公司) ; Florida International University(佛罗里达国际大学) ; University of Illinois - Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 幻觉与鲁棒性 :VLM(title,abstract);vision-language model(abstract);分类 cs.AI
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted by AAAI2026
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV
Comments 16 pages
专题命中 幻觉与鲁棒性 :VLM(title);vision-language model(abstract);分类 cs.CV
Comments we update the paper supplement