arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-30 至 2025-07-30 共收录 25 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 7 篇

2507.21335 2025-07-30 cs.CV 88%

Analyzing the Sensitivity of Vision Language Models in Visual Question Answering

Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal

机构 * University of Memphis, TN, USA(密苏里大学) Adobe Research, San Jose, CA, USA(Adobe研究院) University of Maryland Baltimore County, USA(马里兰大学巴尔的摩县分校)

专题命中 视觉问答 :vision language model(title,abstract);visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18985 2025-07-30 cs.CV cs.AI 84%

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Guanxi Shen

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Keywords: Explainable Computer Vision, Large Vision-Language Models, AI Interpretability, Explainable AI, Visual Saliency, Attribution Maps, Cross-Modal Attribution, Human Attention Alignment, AI Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21585 2025-07-30 cs.AI 77%

SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation

Hao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu, Huadong Ma

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21124 2025-07-30 cs.HC cs.AI cs.GR cs.LG 62%

VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific Visualization

Ayan Biswas, Terece L. Turton, Nishath Rajiv Ranasinghe, Shawn Jones, Bradley Love, William Jones, Aric Hagberg, Han-Wei Shen, Nathan DeBardeleben, Earl Lawrence

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) Coastal Carolina University(海岸卡罗来纳大学) Ohio State University(俄亥俄州立大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09081 2025-07-30 cs.CV cs.AI cs.CL 62%

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Zheqi He, Yesheng Liu, Jing-shu Zheng, Xuejing Li, Jin-Ge Yao, Bowen Qin, Richeng Xuan, Xi Yang

机构 * BAAI FlagEval Team(BAAI 评测团队)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACL 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21165 2025-07-30 eess.IV cs.CV 57%

Querying GI Endoscopy Images: A VQA Approach

Gaurav Parajuli

机构 * Johannes Kepler University Linz(约翰·凯撒大学林茨)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21520 2025-07-30 cs.IR 50%

Solution for Meta KDD Cup'25: A Comprehensive Three-Step Framework for Vision Question Answering

Zijian Zhang, Xiaocheng Zhang, Yang Zhou, Zhimin Lin, Peng Yan

专题命中 视觉问答 :visual question answering(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 1 篇

2507.21917 2025-07-30 cs.CV 70%

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

机构 * Department of Computer Science University of Bari Aldo Moro(计算机科学系巴里大学Aldo Moro)

专题命中 视觉推理 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 7 篇

2507.21507 2025-07-30 cs.CV cs.MM 79%

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding

Shibo Gao, Peipei Yang, Yangyang Liu, Yi Chen, Han Zhu, Xuyao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 21 pages, 19 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21450 2025-07-30 cs.CV cs.RO 79%

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

Bolei Chen, Jiaxu Kang, Yifei Wang, Ping Zhong, Qi Wu, Jianxin Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Submitted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21080 2025-07-30 cs.CL cs.AI 79%

Which symbol grounding problem should we try to solve?

Vincent C. Müller

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref (2015) Journal of Experimental and Theoretical Artificial Intelligence, 27 (1, ed. D. Jones & A. Beavers), 73-78

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19847 2025-07-30 cs.CV 79%

Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection

Wenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng, Lei Zhang

机构 * Hong Kong Polytechnic University(香港理工大学) Eastern Institute of Technology(东部技术研究所) Stanford University(斯坦福大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴堡研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 74%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视觉定位与Grounding :MLLM(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 70%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21353 2025-07-30 cs.CV cs.LG 62%

Group Relative Augmentation for Data Efficient Action Detection

Deep Anil Patel, Iain Melvin, Zachary Izzo, Martin Renqiang Min

机构 * NEC Laboratories America(NEC美洲实验室)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 1 篇

2505.20726 2025-07-30 cs.RO 50%

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

Liu Dai, Haina Wang, Weikang Wan, Hao Su

专题命中 文档图表理解 :vision-language model(abstract)

Comments Project Website: https://manitaskgen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与鲁棒性 7 篇

2403.15836 2025-07-30 cs.CV 88%

VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification

Lanfeng Zhong, Zongyao Huang, Yang Liu, Wenjun Liao, Shichuan Zhang, Guotai Wang, Shaoting Zhang

机构 * School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(电子科技大学机械与电子工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Pathology, Sichuan Clinical Research Center for Cancer, Sichuan Cancer Hospital & Institute, Affiliated Cancer Hospital of University of Electronic Science and Technology of China(pathology department, 四川省癌症临床研究中心, 四川省肿瘤医院及研究所, 电子科技大学附属肿瘤医院) Department of Radiation Oncology, Sichuan Cancer Hospital and Institute, University of Electronic Science and Technology of China(放射肿瘤科, 四川省肿瘤医院及研究所, 电子科技大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV

Comments Accepted at TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21637 2025-07-30 cs.AI 79%

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models

Wanying Wang, Zeyu Ma, Han Zheng, Xin Tan, Mingang Chen

机构 * Shanghai Key Laboratory of Computer Software Testing and Evaluating(上海软件测试与评估 key laboratory) Shanghai Normal University(上海Normal University) TrustAI Pte. Ltd. East China Normal University(东华师范大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.AI

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21521 2025-07-30 cs.CV 79%

Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration

Athmanarayanan Lakshmi Narayanan, Amrutha Machireddy, Ranganath Krishnan

机构 * Intel Labs(英特尔实验室) Intel Corporation(英特尔公司)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

Comments International Joint Conference on Neural Networks 2025 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21494 2025-07-30 cs.LG 79%

Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated Learning

Wenxuan Bao, Ruxi Deng, Ruizhong Qiu, Tianxin Wei, Hanghang Tong, Jingrui He

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.LG

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21794 2025-07-30 cs.CV 74%

Distribution-Based Masked Medical Vision-Language Model Using Structured Reports

Shreyank N Gowda, Ruichi Zhang, Xiao Gu, Ying Weng, Lu Yang

机构 * School of Computer Science, University of Nottingham(计算机科学学院,诺丁汉大学) Department of Computer Science and Technology, School of Informatics, Xiamen University(计算机科学与技术系,信息学院,厦门大学) CHI Lab, University of Oxford(CHI实验室,牛津大学) School of Computer Science, University of Nottingham Ningbo China(计算机科学学院,宁波大学中国)

专题命中 幻觉与鲁棒性 :vision-language model(title);分类 cs.CV

Comments Accepted in MICCAI-W 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 57%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21489 2025-07-30 cs.CV 57%

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

Zhichuan Wang, Yang Zhou, Zhe Liu, Rui Yu, Song Bai, Yulong Wang, Xinwei He, Xiang Bai

机构 * Huazhong Agricultural University(华中农业大学) Shenzhen University(深圳大学) The University of Hong Kong(香港大学) University of Louisville(路易斯安那大学) ByteDance(字节跳动) Huazhong University of Science and Technology(华中科技大学)

专题命中 幻觉与鲁棒性 :MLLM(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. VLM训练与架构 2 篇

2507.13568 2025-07-30 cs.CV 83%

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning

Kaihong Wang, Donghyun Kim, Margrit Betke

机构 * Waymo Boston University(波士顿大学) Korea University(韩国大学)

专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19651 2025-07-30 cs.CV cs.LG cs.PF 81%

PEVLM: Parallel Encoding for Vision-Language Models

Letian Kang, Shixian Luo, Yiqiang Li, Yuxin Yin, Shenxuan Zhou, Xiaoyang Yu, Jin Yang, Yong Wu

机构 * Li Auto Inc.(利自动公司)

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏