arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-01 至 2025-12-01 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 7 篇

2511.19220 2025-12-01 cs.CV cs.AI 90%

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

大视觉语言模型真的在医学图像上具有基础性吗?来自意大利临床视觉问答的证据

Federico Felizzi, Olivia Riccomi, Michele Ferramola, Francesco Andrea Causio, Manuel Del Medico, Vittorio De Vita, Lorenzo De Mori, Alessandra Piscitelli, Pietro Eric Risuleo, Bianca Destro Castaniti, Antonio Cristiano, Alessia Longo, Luigi De Angelis, Mariapia Vassalli, Marcello Di Pumpo

机构 * SIIAM NSBProject Dept. of Life Sciences & Public Health, UCSC(生命科学与公共卫生系,UCSC) ASL RM 4 UCSC Univ. Paris Cité(巴黎Cité大学) Univ. of Pisa(比萨大学)

专题命中 视觉问答 :vision language model(title,abstract);visual question answering(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 研究通过测试四种先进模型在意大利医学问题上的表现,揭示了大视觉语言模型在视觉基础上的差异,强调了临床部署前的严格评估需求。

Comments Accepted at the Workshop on Multimodal Representation Learning for Healthcare (MMRL4H), EurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22341 2025-12-01 cs.CV cs.LG 79%

Unexplored flaws in multiple-choice VQA evaluations

多选式视觉问答评估中的未探索缺陷

Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagodonat, Stephan Günnemann, Leo Schwinn

机构 * Technical University of Munich(慕尼黑技术大学) Volkswagen AG(大众集团) Munich Data Science Institute(慕尼黑数据科学研究所)

专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

AI总结 本文揭示了多选式视觉问答评估中因提示格式变化导致的未探索偏差,指出其对MLLM评估结果的显著影响,并表明现有缓解策略无法有效应对这些新发现的偏见。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22805 2025-12-01 cs.CV cs.LG cs.MM 74%

From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images

从像素到感受:将MLLMs对齐于人类对图像的认知感知

Yiming Chen, Junlin Han, Tianyi Bai, Shengbang Tong, Filippos Kokkinos, Philip Torr

机构 * Oxford University(牛津大学) HKUST(香港科技大学) University College London(伦敦大学学院) New York University(纽约大学)

专题命中 视觉问答 :MLLM(abstract,comments);multimodal large language model(abstract);分类 cs.CV、cs.LG

AI总结 本文提出CogIP-Bench基准,通过后训练提升MLLMs对图像认知属性的对齐能力,并展示其在图像生成中的应用。

Comments Project page with codes/datasets/models: https://follen-cry.github.io/MLLM-Cognition-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18192 2025-12-01 cs.CV cs.AI 73%

ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization

ARIAL:一个用于文档视觉问答的代理框架,具有精确答案定位

Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Dheeraj Kulshrestha, Rajiv Ramnath

机构 * Ohio State University(俄亥俄州立大学) Flairsoft(Flairsoft公司)

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 ARIAL通过代理协调专门工具,实现文档VQA的高精度和可解释性,取得最佳性能和可解释性成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20460 2025-12-01 cs.CV 70%

Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search

关注关键区域:通过自适应缩放搜索实现无需训练的超高清遥感视觉问答

Yunqi Zhou, Chengjie Jiang, Chun Yuan, Jing Li

机构 * Central University of Finance and Economics(中央财经大学) Tsinghua University(清华大学) East China Normal University(华东师范大学)

专题命中 视觉问答 :LLaVA(abstract);visual question answering(abstract);分类 cs.CV

AI总结 ZoomSearch通过自适应缩放搜索实现无需训练的超高清遥感视觉问答,提升准确性和效率。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12099 2025-12-01 cs.CV 57%

TinyRS-R1: Compact Multimodal Language Model for Remote Sensing

TinyRS-R1:用于遥感的紧凑多模态语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中欧技术大学(METU)电子与电气工程系)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

AI总结 TinyRS-R1是一种专为遥感设计的紧凑多模态语言模型,通过四阶段训练实现高效性能,兼具推理增强与低资源消耗。

Comments Accepted to IEEE Geoscience and Remote Sensing Letters (GRSL). Code, models, and the captions for datasets are available at https://github.com/aybora/TinyRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22141 2025-12-01 cs.CL 50%

Bridging the Modality Gap by Similarity Standardization with Pseudo-Positive Samples

通过伪正样本进行相似性标准化以弥合模态差距

Shuhei Yamashita, Daiki Shirafuji, Tatsuhiko Saito

机构 * Mitsubishi Electric Corporation(三菱电机公司)

专题命中 视觉问答 :vision-language model(abstract)

AI总结 通过伪正样本进行相似性标准化,有效弥合跨模态检索中的模态差距,提升检索性能。

Comments Accepted to PACLIC2025

详情

展开后加载摘要…

URL PDF HTML 收藏