Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing China ; School of Cyber Science ; Engineering, Nanjing University of Science ; VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Tianjin China ; Key Laboratory of Ethnic Language Intelligent Analysis ; Security Governance of MOE, Minzu University of China Beijing China ; Institute of Information Engineering, Chinese Academy of Sciences ; School of Cyber Security, University of Chinese Academy of Sciences ; VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University ; Security Governance of MOE, Minzu University of China
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI
Comments Accepted by 2025 ACM MM
专题命中 视觉问答 :visual question answering(abstract)
Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Purdue University(普渡大学) ; The University of Hong Kong(香港大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);LLaVA(abstract);visual reasoning(abstract);MLLM(abstract)
Comments Published at ICCV 2025
机构 * College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院) ; Shanghai Innovation Institute(上海创新研究院) ; Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心) ; School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院) ; Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科) ; MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室) ; Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院) ; Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系) ; Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院) ; X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室) ; Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科) ; Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心) ; Ningbo Institute of Technology, Beihang University(北航宁波理工学院)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI
Comments 6 pages, 3 figures. Accepted, Presented and Published as part of Proceedings of the 6th International Conference on Recent Advantages in Information Technology (RAIT) 2025
Journal ref 2025 6th International Conference on Recent Advances in Information Technology (RAIT), Dhanbad, India, 2025, pp. 1-6
机构 * Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ICCV 2025. arXiv admin note: text overlap with arXiv:2311.06602 by other authors
机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) ; Xiamen Unisound Intelligence Technology Co., Ltd(厦门Unisound智能科技有限公司) ; Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) ; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),中华人民共和国文化和旅游部,中国)
专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV
Comments Accepted by IJCAI 2025
机构 * Lenovo Research(联想研究院) ; School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) ; Institute of Automation, CAS(中国科学院自动化研究所) ; Macau University of Science and Technology(澳门科学理工学院) ; School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV、cs.AI
Comments preprint
机构 * Key Laboratory of Network Data Science and Technology(网络数据科学与技术重点实验室) ; Institute of Computing Technology(计算技术研究所) ; Chinese Academy of Sciences(中国科学院) ; State Key Laboratory of AI Safety(人工智能安全国家重点实验室) ; University of Chinese Academy of Science(中国科学院大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI、cs.LG
Comments Accepted by CCIR25 and published by Springer LNCS or LNAI
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * College of Computer Science, Sichuan University(四川大学计算机学院) ; Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所) ; Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) ; Department of Radiology, West China Hospital, Sichuan University(四川大学华西医院放射科) ; College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院) ; Department of Mathematics, Sichuan University(四川大学数学系) ; School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院) ; Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore and the Singapore Eye Research Institute, Singapore National Eye Centre(新加坡国立大学 Yong Loo Lin 医学院眼科系及新加坡眼科研所、新加坡国家眼科中心)
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
Comments 17 pages, 6 figures
机构 * Xi’an Jiaotong University(西安交通大学) ; ARC Lab, Tencent PCG(腾讯PCG ARC实验室) ; City University of Hongkong(香港城市大学) ; Institute of Automation, CAS(中国科学院自动化研究所) ; Harvard University(哈佛大学) ; vivo Mobile Communication Co.(vivo移动通信公司)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
Comments 22 pages, 16 figures
机构 * Technical University of Munich(慕尼黑技术大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; Imperial College London(伦敦帝国学院) ; University of Trento(特伦托大学) ; Helmholtz AI and Helmholtz Munich(海德堡人工智能与海德堡慕尼黑) ; King’s College London(伦敦国王学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) ; University of Minnesota(明尼苏达大学) ; Cisco Research(思科研究)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.AI
Comments 9 pages
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王宣计算机技术研究所) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) ; Central Media Technology Institute, Huawei(华为中央媒体技术研究所)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Accepted by ICCV 2025
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * Laboratory of IEMN, Univ. Polytechnique Hauts-de-France(IEMN实验室,法国高等技术法国大学) ; KU 6G Research Center, Khalifa University(KU 6G研究中心,哈利法大学) ; Sorbonne Center for Artificial Intelligence, Sorbonne University(人工智能研究中心,索邦大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments Project page: https://github.com/jasongief/TGS-Agent
机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) ; Shenzhen University of Advanced Technology(深圳大学先进技术学院)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract)
Comments Position paper on LLM-agent simulation of financial structures. This update clarifies setup and adds a reciprocity-based table. Builds on arXiv:2505.02945 and 2505.08319
专题命中 视觉定位与Grounding :VLM(abstract)
专题命中 视觉定位与Grounding :grounding(abstract)
机构 * Microsoft Research(微软研究院)
专题命中 视觉定位与Grounding :grounding(abstract)
Comments First three authors contributed equally (ordered alphabetically)
专题命中 视觉定位与Grounding :grounding(abstract)
Journal ref The Art, Science, and Engineering of Programming, 2025, Vol. 10, Issue 2, Article 17
机构 * independent researcher(独立研究者)
专题命中 文档图表理解 :multimodal large language model(abstract);分类 cs.CV、cs.AI
机构 * George Bredis(无) ; Stanislav Dereka(无) ; Viacheslav Sinii(无) ; Ruslan Rakhimov(无) ; Daniil Gavrilov(无)
专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract);分类 cs.AI、cs.LG
机构 * Zhejiang University(浙江大学) ; Fudan University(复旦大学) ; OPPO AI Center(OPPO AI中心) ; University of Chinese Academy of Sciences(中国科学院大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; The Chinese University of Hong Kong(香港中文大学) ; Tsinghua University(清华大学) ; Shanghai Jiao Tong University(上海交通大学) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 GUI与屏幕智能体 :MLLM(title);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
Comments ACL 2025 (Oral)
机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈格浦尔分校) ; IIIS, Tsinghua University(清华大学人工智能研究所) ; Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG
Comments 7 pages. Accepted at the LangRob Workshop 2024 @ CoRL, 2024. Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)