arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-26 至 2025-08-26 共收录 57 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 6 篇

2508.16661 2025-08-26 cs.CV 88%

QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models

Qiaojie Zheng, Jiucai Zhang, Joy Gockel, Michael B. Wakin, Craig Brice, Xiaoli Zhang

机构 * Mechanical Engineering, Colorado School of Mines(机械工程,科罗拉多矿业学院) Electrical Engineering, Colorado School of Mines(电气工程,科罗拉多矿业学院)

专题命中 视觉问答 :VLM(title,abstract);vision language model(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17860 2025-08-26 cs.CV cs.AI 84%

AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering

Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan, Zhiyong Li

机构 * Hunan University(湖南大学) South China Normal University(华南师范大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16763 2025-08-26 cs.CV 77%

WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow Mila Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) École de Technologie Supérieure (ETS)(高等技术学院) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments This paper has been accepted to the EMNLP 2025 main conference. Check the project page here: https://webmmu-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15075 2025-08-26 cs.CL cs.AI cs.CV cs.LG 75%

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie, Daisuke Kawahara

专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments The first version of this paper mistakenly included a prompt injection phrase, which was inappropriate and unprofessional. Although we corrected the version on arXiv and withdrew from the conference, my co-authors and university strongly request a full withdrawal. Given the situation, I no longer have the authority to manage this paper, and withdrawing it from arXiv is the most responsible action

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17398 2025-08-26 cs.CL 50%

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

Aaryaman Kartha, Ahmed Masry, Mohammed Saidul Islam, Thinh Lang, Shadikur Rahman, Ridwan Mahbub, Mizanur Rahman, Mahir Ahmed, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty

机构 * York University(约克大学) RBC Qatar Computing Research Institute (QCRI)(卡塔尔计算研究所) Nanyang Technological University(南洋理工大学) Salesforce Research(Salesforce研究)

专题命中 视觉问答 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18023 2025-08-26 cs.CL 50%

Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference

Zhuo Chen, Xinyu Wang, Yong Jiang, Zhen Zhang, Xinyu Geng, Pengjun Xie, Fei Huang, Kewei Tu

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Shanghai Engineering Research Center of Intelligent Vision and Imaging(智能视觉与成像上海工程研究中心) Institute for Intelligent Computing, Alibaba Group(智能计算研究院,阿里巴巴集团)

专题命中 视觉问答 :visual question answering(abstract)

Comments EMNLP2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 13 篇

2508.18227 2025-08-26 cs.CV 85%

GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models

Lianming Huang, Haibo Hu, Qiao Li, Xin He, Nan Guan, Chun Jason Xue

机构 * City University of Hong Kong(香港城市大学) MBZUAI A*STAR

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15969 2025-08-26 cs.CV cs.AI cs.CL 84%

Forgotten Polygons: Multimodal Large Language Models are Shape-Blind

William Rudman, Michal Golovanevsky, Amir Bar, Vedant Palit, Yann LeCun, Carsten Eickhoff, Ritambhara Singh

机构 * Brown University(布朗大学) Tel Aviv University(特拉维夫大学) IIT Kharagpur(印度理工学院卡里帕尔分校) New York University(纽约大学) University of Tübingen(图宾根大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18179 2025-08-26 cs.AI cs.CV 81%

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

Zhenwei Tang, Difan Jiao, Blair Yang, Ashton Anderson

机构 * Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Coolwei AI Lab(Coolwei人工智能实验室)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17290 2025-08-26 cs.AI cs.LG 73%

MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment

Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah

机构 * Computer Engineering Department, Sharif University of Technology, Iran(谢尔盖大学计算机工程系,伊朗) Qatar Computing Research Institute, Qatar(卡塔尔计算研究所,卡塔尔) Computer Engineering Department, University of Isfahan, Iran(伊斯法罕大学计算机工程系,伊朗) Independent Researcher(独立研究者)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17205 2025-08-26 cs.CV cs.AI cs.CL eess.IV 73%

Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding

Yunxiang Yang, Ningning Xu, Jidong J. Yang

机构 * Smart Mobility and Infrastructure Lab(智能移动与基础设施实验室) College of Engineering, University of Georgia(佐治亚大学工程学院)

专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 16 pages, 16 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12687 2025-08-26 cs.AI cs.CV 73%

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Virginia(弗吉尼亚大学)

专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08189 2025-08-26 cs.LG cs.AI 73%

DSADF: Thinking Fast and Slow for Decision Making

Zhihao Dou, Dongfei Cui, Jun Yan, Weida Wang, Benteng Chen, Haoming Wang, Zeke Xie, Shufei Zhang

机构 * Pratt School of Engineering(普拉特工程学院) Duke University(杜克大学) School of Computer Science(计算机科学学院) Northeast Electric Power University(东北电力大学) Department of Information and Communication Engineering(信息与通信工程系) Tongji University(同济大学) School of Computer Science and Technology(计算机科学与技术学院) School of International Education(国际教育学院) Beijing University of Chemical Technology(北京化工大学) East China Normal University(华东师范大学) Information Hub(信息枢纽) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Science(中国科学院)

专题命中 视觉推理 :vision language model(abstract);VLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17595 2025-08-26 cs.CV 57%

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints

Vinh-Thuan Ly, Hoang M. Truong, Xuan-Huong Nguyen

机构 * University of Science, VNU-HCM(越南胡志明市国家大学) Vietnam National University(越南国家大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

Comments Accepted for presentation at the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2025

Journal ref IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Hawaii, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08276 2025-08-26 cs.CL cs.AI 57%

Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models

Yassine Jamaa, Badr AlKhamissi, Satrajit Ghosh, Martin Schrimpf

机构 * EPFL(苏黎世联邦理工学院) MIT(麻省理工学院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI

Comments Accepted at the Interplay of Model Behavior and Model Internals Workshop co-located with COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17589 2025-08-26 cs.AI 57%

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown

Bowen Wang, Zhouqiang Jiang, Yasuaki Susumu, Shotaro Miwa, Tianwei Chen, Yuta Nakashima

机构 * Osaka University(大阪大学) Mitsubishi Electric Corp.(三菱电机公司)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI

Comments Aligned with ICCV 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00258 2025-08-26 cs.AI 57%

Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs

Qianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao, Xinze Guan, Ching-Chen Kuo, Xin Eric Wang

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) eBay

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16859 2025-08-26 cs.CV 57%

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li, Peipei Song, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (IAI), Hefei Comprehensive National Science Center(人工智能研究院(IAI),合肥综合性国家科学中心)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16850 2025-08-26 cs.AI 57%

RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis

Anku Rani, Aparna Garimella, Apoorv Saxena, Balaji Vasan Srinivasan, Paul Pu Liang

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 18 篇

2505.15123 2025-08-26 cs.CV cs.AI 84%

Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

Ta Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi, Qi Chen, Phi Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan W. Verjans, Vu Minh Hieu Phan

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Macquarie University(麦考瑞大学) Hanoi University of Science and Technology(河内科学技术大学) University of Wollongong(沃林根大学) Flinders University(弗林德斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17976 2025-08-26 cs.CV eess.IV 83%

Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization

Keyang Zhang, Chenqi Kong, Hui Liu, Bo Ding, Xinghao Jiang, Haoliang Li

机构 * Department of Electrical Engineering, City University of Hong Kong(香港城市大学电子工程系) Rapid-Rich Object Search (ROSE) Lab, School of Electrical and Electronic Engineering, Nanyang Technology University(南洋理工大学电子与电气工程学院快速丰富对象搜索(ROSE)实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :MLLM(title);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16974 2025-08-26 cs.CV 83%

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren, Zheyu Song

机构 * Zhongkai University of Agriculture and Engineering(仲恺农业工程大学) EDP University of Puerto Rico: San Sebastian(波多黎各圣塞巴斯蒂安EDP大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02943 2025-08-26 cs.CR cs.AI cs.CL cs.LG 81%

PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding

Krishna Kanth Nakka, Ahmed Frikha, Ricardo Mendes, Xue Jiang, Xuebing Zhou

机构 * Huawei Munich Research Center(华为慕尼黑研究中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments Accepted at PrivateNLP Workshop at ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 79%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18132 2025-08-26 cs.IR cs.AI cs.LG 73%

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang

机构 * Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以) NVIDIA(NVIDIA公司) Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13357 2025-08-26 cs.CL 71%

Adaptive Linguistic Prompting (ALP) Enhances Phishing Webpage Detection in Multimodal Large Language Models

Atharva Bhargude, Ishan Gonehal, Dave Yoon, Kaustubh Vinnakota, Chandler Haney, Aaron Sandoval, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 视觉定位与Grounding :multimodal large language model(title)

Comments Published at ACL 2025 SRW, 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02259 2025-08-26 cs.CV 70%

T*: Re-thinking Temporal Search for Long-Form Video Understanding

Jinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristobal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli, Li Fei-Fei, Jiajun Wu, Manling Li

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);分类 cs.CV

Comments Accepted by CVPR 2025; A real-world long video needle-in-haystack benchmark; long-video QA with human ref frames

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18742 2025-08-26 cs.CV cs.RO 70%

3D Feature Distillation with Object-Centric Priors

Georgios Tziafas, Yucheng Xu, Zhibin Li, Hamidreza Kasaei

机构 * Department of Artificial Intelligence University of Groningen, the Neteherlands(格罗宁根大学人工智能系) School of Informatics University of Edinburgh, United Kingdom(爱丁堡大学信息学院) Department of Computer Science University College London, United Kingdom(伦敦大学学院计算机科学系)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17667 2025-08-26 cs.CV cs.AI 62%

Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection

Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Nottingham Malaysia(诺丁汉大学(马来西亚)) Key Laboratory of Machine Intelligence and Advanced Computing, MOE(机器智能与高级计算重点实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 2 figures, Accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16643 2025-08-26 cs.LG cs.AI 62%

From Classical Probabilistic Latent Variable Models to Modern Generative AI: A Unified Perspective

Tianhua Chen

机构 * School of Computing and Engineering University of Huddersfield(计算与工程学院赫德斯菲尔德大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments This is a substantially improved and expanded version of an earlier manuscript hosted on SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5244929

详情

展开后加载摘要…

URL PDF HTML 收藏