arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-21 至 2025-10-21 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2510.16292 2025-10-21 cs.LG 85%

QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models

Yutong Wang, Haiyu Wang, Sai Qian Zhang

机构 * Tandon School of Engineering, New York University(纽约大学工程学院) Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)

专题命中 视觉问答 :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.LG

Comments Accepted as Spotlight paper by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17771 2025-10-21 cs.AI cs.CV 85%

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, Benoit Dumoulin, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊) Penn State University(宾夕法尼亚州立大学)

专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);LLaVA(abstract);InternVL(abstract)

Comments 21 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14605 2025-10-21 cs.CV cs.AI 73%

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Alibaba Cloud Computing(阿里巴巴云计算)

专题命中 视觉问答 :visual language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13946 2025-10-21 cs.AI 70%

Visual Instruction Bottleneck Tuning

Changdae Oh, Jiatong Li, Shawn Im, Sharon Li

机构 * Department of Computer Sciences, University of Wisconsin–Madison(计算机科学系,威斯康星大学麦迪逊分校)

专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏