When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract)
Comments Accepted by AAAI 2026
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract)
Comments Accepted by AAAI 2026
机构 * EXL Health AI Lab at MEDIQA-WV 2025(EXL健康AI实验室)
专题命中 视觉问答 :visual question answering(title);分类 cs.AI
Comments 2 figures, 11 pages
机构 * KAUST(卡斯泰尔大学)
专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG
Comments 12 pages, 7 figures, AI4NextG @ NeurIPS 2025
专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV
Comments Accepted in AAAI 2026
专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV
Comments Accepted by AAAI 2026
机构 * University of Melbourne(墨尔本大学)
专题命中 视觉推理 :visual reasoning(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * University of Science and Technology of China(科学技术大学) ; Singapore University of Technology and Design(新加坡科技设计大学) ; Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中 视觉推理 :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) ; School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) ; Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) ; School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) ; Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))
专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract)
Comments 12 pages, 9 figures
机构 * Integrated Program in Neuroscience (IPN) McGill University(神经科学联合计划 麦吉尔大学) ; Mila, University of Montreal(蒙特利尔大学Mila) ; Microsoft Research USA(微软研究院美国总部) ; Department of Physiology McGill University(生理学系 麦吉尔大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
Comments Paper accepted by nips 2025
机构 * Nanjing University(南京大学) ; CASIA(中国科学院自动化研究所) ; Kuaishou Technology(快手科技) ; M-A-P
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI
Journal ref The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)
机构 * Xi’an Jiaotong University(西安交通大学) ; University of Science and Technology of China(中国科学技术大学) ; SenseTime Research(商汤科技研究院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
Comments [Accepted by AAAI2026] Project Page: https://zju-real.github.io/gui-rcpo Code: https://github.com/zju-real/gui-rcpo
专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI
机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) ; Bosch Global Software Technologies(博世全球软件技术公司)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments Accepted to AAAI 2026 (Oral), Project Page: https://github.com/JiuTian-VL/SemanticVLA
机构 * City University of Hong Kong(香港城市大学) ; Monash University(墨尔本大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments Project Page: https://layerpeeler.github.io/
专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract)
Comments Accepted by AAAI 2026 (Oral)
机构 * Ketong Chen, Yuhao Chen, Yang Xue(作者)
专题命中 文档图表理解 :vision-language model(abstract);分类 cs.CV
机构 * Department of Computing, University of Turku(图尔库大学计算机系) ; Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所) ; Centre for Artificial Ingelligence, Zurich University of Applied Sciences(应用科学大学人工智能中心) ; Agentic Systems Lab, Department of Management, Technology and Economics, ETH Zürich(苏黎世联邦理工学院管理、科技与经济系代理系统实验室) ; Faculty of Mathematics and Information Science, Warsaw University of Technology(华沙技术大学数学与信息科学学院)
专题命中 GUI与屏幕智能体 :VLM(title);vision-language model(abstract);分类 cs.AI、cs.LG
机构 * Indian Institute of Technology Mandi(印度理工学院曼迪分校) ; Vellore Institute of Technology(韦洛雷理工学院) ; Indian Institute of Technology Kharagpur(印度理工学院哈里科普分校)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV
Comments 8 pages
专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments To appear in the AI4NextG Workshop at NeurIPS 2025
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI
Comments accepted for publication in the Association for the Advancement of Artificial Intelligence (AAAI), 2026
机构 * Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系) ; Translational Data Analytics Institute, The Ohio State University(俄亥俄州立大学转化数据分析研究所) ; Department of Biomedical Informatics, The Ohio State University(俄亥俄州立大学生物医学信息学系)
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI
Comments ICJNLP-AACL 2025
机构 * Amazon Bedrock Science(亚马逊Bedrock科学) ; Drexel University(德雷塞尔大学) ; University of Virginia(弗吉尼亚大学)
专题命中 幻觉与鲁棒性 :VLM(abstract);分类 cs.AI
Comments 14 pages, 5 figures; published in EMNLP 2025 ; Code at: https://github.com/dsbuddy/GAP-LLM-Safety
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract)
Comments Accepted at AAAI 2026
专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by AAAI 2026
专题命中 VLM训练与架构 :vision-language model(title);visual language model(abstract);分类 cs.CV
Comments AAAI2026, with supplementary material
专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV