arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-26 至 2025-08-26 共收录 6 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 6 篇

2508.16661 2025-08-26 cs.CV 88%

QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models

Qiaojie Zheng, Jiucai Zhang, Joy Gockel, Michael B. Wakin, Craig Brice, Xiaoli Zhang

机构 * Mechanical Engineering, Colorado School of Mines(机械工程,科罗拉多矿业学院) Electrical Engineering, Colorado School of Mines(电气工程,科罗拉多矿业学院)

专题命中 视觉问答 :VLM(title,abstract);vision language model(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17860 2025-08-26 cs.CV cs.AI 84%

AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering

Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan, Zhiyong Li

机构 * Hunan University(湖南大学) South China Normal University(华南师范大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16763 2025-08-26 cs.CV 77%

WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow Mila Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) École de Technologie Supérieure (ETS)(高等技术学院) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments This paper has been accepted to the EMNLP 2025 main conference. Check the project page here: https://webmmu-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15075 2025-08-26 cs.CL cs.AI cs.CV cs.LG 75%

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie, Daisuke Kawahara

专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments The first version of this paper mistakenly included a prompt injection phrase, which was inappropriate and unprofessional. Although we corrected the version on arXiv and withdrew from the conference, my co-authors and university strongly request a full withdrawal. Given the situation, I no longer have the authority to manage this paper, and withdrawing it from arXiv is the most responsible action

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17398 2025-08-26 cs.CL 50%

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

Aaryaman Kartha, Ahmed Masry, Mohammed Saidul Islam, Thinh Lang, Shadikur Rahman, Ridwan Mahbub, Mizanur Rahman, Mahir Ahmed, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty

机构 * York University(约克大学) RBC Qatar Computing Research Institute (QCRI)(卡塔尔计算研究所) Nanyang Technological University(南洋理工大学) Salesforce Research(Salesforce研究)

专题命中 视觉问答 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18023 2025-08-26 cs.CL 50%

Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference

Zhuo Chen, Xinyu Wang, Yong Jiang, Zhen Zhang, Xinyu Geng, Pengjun Xie, Fei Huang, Kewei Tu

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Shanghai Engineering Research Center of Intelligent Vision and Imaging(智能视觉与成像上海工程研究中心) Institute for Intelligent Computing, Alibaba Group(智能计算研究院,阿里巴巴集团)

专题命中 视觉问答 :visual question answering(abstract)

Comments EMNLP2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏