arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3138 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3138 篇

2412.04939 2025-12-23 cs.CV 79%

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

动词幻觉:揭示和评估多模态大语言模型中的动词概念幻觉

Zehao Wang, Xinpeng Liu, Yudonglin Zhang, Xiaoqian Wu, Zhou Fang, Yifan Fang, Junfu Pu, Cewu Lu, Yong-Lu Li

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

AI总结 本文首次揭示多模态大语言模型中的动词概念幻觉问题,并提出基于丰富动词知识的调优方法以有效缓解该问题。

Comments Accepted by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12799 2025-12-16 cs.CV 79%

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

DrivePI: 基于空间感知的4D MLLM用于统一自动驾驶理解、感知、预测与规划

Zhe Liu, Runhui Huang, Rui Yang, Siming Yan, Zining Wang, Lu Hou, Di Lin, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Yinwang Intelligent Technology Co. Ltd.(英维智能科技有限公司) Tianjin University(天津大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 视觉问答 :MLLM(title,abstract);分类 cs.CV

AI总结 DrivePI是一种基于空间感知的4D MLLM,用于统一自动驾驶的理解、感知、预测和规划,通过端到端优化实现多任务并行,提升性能并减少碰撞率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11510 2025-12-15 cs.CV 79%

Reconstruction as a Bridge for Event-Based Visual Question Answering

重建作为事件驱动视觉问答的桥梁

Hanyue Lou, Jiayi Zhou, Yang Zhang, Boyu Li, Yi Wang, Guangnan Ye, Boxin Shi

机构 * Peking University(北京大学) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学)

专题命中 视觉问答 :visual question answering(title);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出基于重建的框架,通过FRT和ART方法提升事件驱动视觉问答性能,并引入EvQA基准验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04686 2025-12-09 cs.CV 79%

Towards Cross-View Point Correspondence in Vision-Language Models

面向视觉-语言模型的跨视图点对应研究

Yipu Wang, Yuheng Ji, Yuyang Liu, Enshen Zhou, Ziqiang Yang, Yuxuan Tian, Ziheng Qin, Yue Liu, Huajie Tan, Cheng Chi, Zhiyuan Ma, Daniel Dajun Zeng, Xiaolong Zheng

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Beihang University(北航大学) Jilin University(吉林大学) National University of Singapore(新加坡国立大学) Peking University(北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Huazhong University of Science And Technology(华中科技大学)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出跨视图点对应任务和 CrossPoint-Bench 基准,通过 CroPond 模型在 CrossPoint-Bench 上取得超越 Gemini-2.5-Pro 的准确率,推动跨视图对应研究发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22154 2025-12-03 cs.AI 79%

WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios

WearVQA: 一个用于评估智能眼镜等可穿戴设备上多模型AI助手视觉问答能力的基准测试

Eun Chang, Zhuangqun Huang, Yiwei Liao, Sagar Ravi Bhavsar, Amogh Param, Tammy Stark, Adel Ahmadyan, Xiao Yang, Jiaqi Wang, Ahsan Abdullah, Giang Nguyen, Akil Iyer, David Hall, Elissa Li, Shane Moon, Nicolas Scheffer, Kirmani Ahmed, Babak Damavandi, Rakesh Wanga, Anuj Kumar, Rohit Patel, Xin Luna Dong

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

AI总结 WearVQA是一个评估可穿戴设备上多模型AI助手视觉问答能力的基准测试,通过真实场景下的图像-问题-答案三元组,评估其在复杂视觉输入和现实任务中的表现。

Comments 11 pages, 5 figures, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00773 2025-12-02 cs.CV 79%

DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering

DEJIMA:一种新的大规模日语数据集用于图像描述和视觉问答

Toshiki Katsube, Taiga Fukuhara, Kenichiro Ando, Yusuke Mukuta, Kohei Uehara, Tatsuya Harada

机构 * The University of Tokyo(东京大学) RIKEN(日本资源技术研究所)

专题命中 视觉问答 :visual question answering(title);grounding(abstract);分类 cs.CV

AI总结 DEJIMA是首个大规模日语图像描述与视觉问答数据集,通过整合网络收集、目标检测和大语言模型优化,提升了日语语言和文化表现力,显著优于现有翻译或人工标注数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14567 2025-11-24 cs.HC cs.AI 79%

SweeperBot: Making 3D Browsing Accessible through View Analysis and Visual Question Answering

SweeperBot:通过视图分析和视觉问答实现3D浏览的可访问性

Chen Chen, Cuong Nguyen, Alexa Siu, Dingzeyu Li, Nadir Weibel

机构 * Florida International University(佛罗里达国际大学) Adobe Research(Adobe研究) University of California San Diego(加州大学圣地亚哥分校)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

AI总结 SweeperBot通过视图分析和视觉问答技术,帮助屏幕阅读器用户更有效地探索和比较3D模型。

Comments 28 pages, 16 figures, this is an original manuscript of an article published by Taylor & Francis in the International Journal of Human-Computer Interaction (IJHCI), available online: https://doi.org/10.1080/10447318.2025.2594750

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11198 2025-11-17 cs.CV 79%

Geospatial Chain of Thought Reasoning for Enhanced Visual Question Answering on Satellite Imagery

Shambhavi Shanker, Manikandan Padmanaban, Jagabondhu Hazra

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔) IBM Research India(IBM印度研究)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09058 2025-11-13 cs.CV 79%

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology, VNU University of Engineering and Technology(信息科技学院,越南工程与技术大学) IT-BT Convergence Technology Division, Vietnam-Korea Institute of Science and Technology(IT-BT融合技术部,越南-韩国科学技术院) TADI Global Lab, TADI Global Company Limited(TADI全球实验室,TADI全球公司) Faculty of Finance, Banking Academy of Vietnam(金融学院,越南银行学院)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures, 3 tables, FAIR 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03617 2025-11-06 cs.GR cs.AI 79%

Visualization Biases MLLM's Decision Making in Network Data Tasks

Timo Brand, Henry Förster, Stephen G. Kobourov, Jacob Miller

机构 * Technical University of Munich, Heilbronn, Germany(慕尼黑技术大学)

专题命中 视觉问答 :MLLM(title,abstract);分类 cs.AI

Comments This manuscript was presented at VIS x GenAI, a workshop co-located with IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00389 2025-11-04 cs.CV 79%

Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond

Fan Zhang, Haoxuan Li, Shengju Qian, Xin Wang, Zheng Lian, Hao Wu, Zhihong Zhu, Yuan Gao, Qiankun Li, Yefeng Zheng, Zhouchen Lin, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学) Tencent(腾讯) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22803 2025-10-28 cs.CV 79%

MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology VNU University of Engineering(信息科技学院越南工程大学) IT-BT Convergence Technology Division Vietnam-Korea Institute of Science(IT-BT融合技术部门越南-韩国科学技术院) TADI Global Lab TADI Global Company Limited(TADI全球实验室TADI全球有限公司) Faculty of Finance Banking Academy of Vietnam(金融学院越南银行学院)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures, IEEE conference format

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08974 2025-10-27 cs.CV 79%

Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering

Elman Ghazaei, Erchan Aptoula

机构 * Faculty of Engineering and Natural Sciences (VPALab)(工程与自然科学学院(VPALab)) Sabanci University(萨班奇大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07447 2025-10-15 cs.CV 79%

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, School of CSE, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学计算机科学与工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 视觉问答 :MLLM(title);multimodal large language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11645 2025-10-14 cs.AI 79%

Adapting and Evaluating Multimodal Large Language Models for Adolescent Idiopathic Scoliosis Self-Management: A Divide and Conquer Framework

Zhaolong Wu, Pu Luo, Nan Meng, Jason Pui Yin Cheung, Teng Zhang

机构 * Department of Orthopaedics and Traumatology, The University of Hong Kong(骨科与创伤外科部,香港大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.AI

Comments Accepted by MICCAI 2025 MLLMCP Workshop

Journal ref Lecture Notes in Computer Science 16147 (2026) 280-289

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08791 2025-10-13 cs.CV 79%

Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering

Yuanhao Zou, Zhaozheng Yin

机构 * University of Michigan(密歇根大学) Stony Brook University(石溪大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments CVPR2025 Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03903 2025-10-07 cs.CV 79%

Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models

Md. Atabuzzaman, Andrew Zhang, Chris Thomas

机构 * Department of Computer Science(计算机科学系) Virginia Tech(弗吉尼亚理工大学)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03232 2025-10-06 cs.CV 79%

LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models

Ci-Siang Lin, Min-Hung Chen, Yu-Yang Sheng, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾国立台湾大学通信工程研究所) NVIDIA

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23899 2025-10-06 cs.CV 79%

Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq, Rakesh Chandra Joshi

机构 * Amity University(阿米蒂大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 12 pages (9 main + 2 references/appendix), 2 figures, conference paper submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17974 2025-09-23 cs.CL cs.CV 79%

Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts

Xuyang Wu, Yuan Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang

机构 * Santa Clara University(圣克拉拉大学) DOCOMO Innovations, Inc.(DOCOMO创新公司) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14886 2025-09-19 cs.CL cs.AI 79%

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Ye Shen, Junying Wang, Farong Wen, Yijin Guo, Qi Jia, Zicheng Zhang, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学)

专题命中 视觉问答 :MLLM(title,abstract);分类 cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06010 2025-09-09 cs.CV 79%

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Wanyin Cheng, Zanxi Ruan

机构 * School of Cyber Science and Engineering, Qufu Normal University(网络科学与工程学院,曲阜师范大学) Department of Computer Science, University of Verona(计算机科学系,威尼斯大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04502 2025-09-08 cs.CL cs.AI 79%

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

Qixin Sun, Ziqin Wang, Hengyuan Zhao, Yilin Li, Kaiyou Song, Linjiang Huang, Xiaolin Hu, Qingpei Guo, Si Liu

专题命中 视觉问答 :multimodal large language model(title);visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19887 2025-08-28 cs.CL cs.CV 79%

Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement

Mohammed Rakibul Hasan, Rafi Majid, Ahanaf Tahmid

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12404 2025-08-19 cs.CV 79%

LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving

Nan Song, Bozhou Zhang, Xiatian Zhu, Jiankang Deng, Li Zhang

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05262 2025-08-15 cs.CV 79%

Debiasing Multimodal Large Language Models via Penalization of Language Priors

YiFan Zhang, Yang Shi, Weichen Yu, Qingsong Wen, Xue Wang, Wenjing Yang, Zhang Zhang, Liang Wang, Rong Jin

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学) Alibaba Group(阿里巴巴集团) National University of Defense Technology(国防科技大学) Meta

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

Comments 10 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00549 2025-08-11 cs.CV 79%

Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images

Daniel Wolf, Heiko Hillenhagen, Billurvan Taskin, Alex Bäuerle, Meinrad Beer, Michael Götz, Timo Ropinski

机构 * Visual Computing Group, Institute of Media Informatics, Ulm University, Germany(媒体信息研究所视觉计算组,乌尔姆大学,德国) Diagnostic and Interventional Radiology, Ulm University Medical Center, Germany(乌尔姆大学医学中心诊断与介入放射学) Axiom Bio, USA(Axiom Bio公司,美国)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18351 2025-08-08 cs.CL cs.AI 79%

Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Zhenqi Wang

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16936 2025-08-08 cs.CL cs.AI 79%

Rationale-guided Prompting for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Fengyuan Liu

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17050 2025-07-24 cs.CV 79%

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas, Subarna Tripathi

机构 * Intel(英特尔)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted to CVAM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏