Questioning the Stability of Visual Question Answering
机构 * Bar-Ilan University(巴伊兰大学)
专题命中 视觉问答 :visual question answering(title);VLM(abstract);visual language model(abstract);分类 cs.CV、cs.LG
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Bar-Ilan University(巴伊兰大学)
专题命中 视觉问答 :visual question answering(title);VLM(abstract);visual language model(abstract);分类 cs.CV、cs.LG
机构 * College of Computer, National University of Defense Technology(国防科技大学计算机学院) ; School of Computer and Information Engineering, Hefei University of Technology(合肥工业大学计算机与信息工程学院) ; CCNU(中国地质大学)
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
专题命中 视觉问答 :grounding(title);vision-language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉问答 :vision-language model(abstract);LLaVA(abstract);visual question answering(abstract);multimodal large language model(abstract)
机构 * TCS Research(塔塔咨询研究)
专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Comments 17 pages, 6 figures, 5 tables. Accepted to Special Track on AI Alignment, AAAI 2026. Project Page- https://refine-align.github.io/
机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔) ; IBM Research India(IBM印度研究)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :grounding(abstract);multimodal large language model(abstract);分类 cs.AI
Comments Accepted by AAAI 2026 Artificial Intelligence for Social Impact Track
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI
机构 * University of Southern California(南加州大学)
专题命中 视觉推理 :vision-language model(title,abstract);visual reasoning(abstract);分类 cs.CV
机构 * Fudan University(复旦大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
Comments 11 pages, 4 figures
机构 * PrismaX Team Shanghai Artificial Intelligence Laboratory(普斯玛X团队上海人工智能实验室)
专题命中 视觉推理 :MLLM(title);InternVL(abstract);multimodal large language model(abstract);分类 cs.AI
Comments 82 pages
专题命中 视觉推理 :multimodal large language model(title,abstract)
Comments 9 pages, 5 figures
机构 * National University of Singapore(新加坡国立大学) ; Tencent Youtu Lab(腾讯优图实验室) ; Tsinghua University(清华大学) ; University of Science and Technology of China(中国科学技术大学) ; Nanyang Technological University(南洋理工大学) ; Zhejiang University(浙江大学) ; DeepWisdom(深智科技)
专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
机构 * The Alan Turing Institute(艾伦·图灵研究所)
专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.AI
机构 * National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学) ; University of Maryland, College Park(马里兰大学学院公园分校) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :VLM(abstract);分类 cs.CV
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments 15 pages, 5 figures
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV
Comments This is the extended version of the paper accepted at AAAI 2026, which includes all technical appendices and additional experimental details
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
Comments Accepted by AAAI 2026
机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments Accepted to IROS 2025
机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) ; University of Florida(佛罗里达大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
机构 * German Cancer Research Center, Division of Medical Image Computing(德国癌症研究中心,医学影像计算部) ; Faculty of Mathematics and Computer Science(数学与计算机科学系) ; Medical Faculty - Heidelberg University(海德堡大学医学系) ; Helmholtz Imaging(海德堡影像技术) ; Department of Radiation Oncology, Heidelberg University Hospital(海德堡大学医院放射肿瘤科) ; HIDSS4Health, Heidelberg(HIDSS4Health,海德堡) ; Pattern Analysis and Learning Group, Heidelberg University Hospital(海德堡大学医院模式分析与学习组)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG
机构 * Mohamed bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)
专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV
Comments 2 figures, 3 tables
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) ; School of Mathematical Sciences, Fudan University(复旦大学数学学院) ; School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络科学与技术学院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
机构 * Google Cloud AI Research(谷歌云人工智能研究) ; School of Computer Science, Peking University(北京大学计算机学院)
专题命中 文档图表理解 :vision-language model(abstract);分类 cs.CV
机构 * Deakin University(德克萨斯大学) ; Sungkyunkwan University(成均馆大学)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV
Comments 12 pages, Under review
专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.CV
Comments 23 pages
机构 * Department of Data and Decision Science, Technion - Israel Institute of Technology(数据与决策科学系,技术学院-以色列理工学院) ; Faculty of Computer and Information Science, Ben-Gurion University of the Negev(计算机与信息科学系,贝内-约尔根大学)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.LG
Comments Accepted to The First Workshop on Confabulation, Hallucinations, & Overgeneration in Multilingual & Precision-critical Setting - AACL-IJCNLP2025
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV、cs.AI
Comments AAAI 2026. Code: https://github.com/HaokunChen245/AUVIC