arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-11 至 2025-08-11 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2508.00549 2025-08-11 cs.CV 79%

Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images

Daniel Wolf, Heiko Hillenhagen, Billurvan Taskin, Alex Bäuerle, Meinrad Beer, Michael Götz, Timo Ropinski

机构 * Visual Computing Group, Institute of Media Informatics, Ulm University, Germany(媒体信息研究所视觉计算组,乌尔姆大学,德国) Diagnostic and Interventional Radiology, Ulm University Medical Center, Germany(乌尔姆大学医学中心诊断与介入放射学) Axiom Bio, USA(Axiom Bio公司,美国)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06152 2025-08-11 cs.CV 57%

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

Kaiyuan Jiang, Ruoxi Sun, Ying Cao, Yuqi Xu, Xinran Zhang, Junyan Guo, ChengSheng Deng

机构 * Peking University(北京大学) LinkSure University of Glasgow(格拉斯哥大学) Boston University(波士顿大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments 17 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08131 2025-08-11 cs.CV 57%

SAR Strikes Back: A New Hope for RSVQA

Lucrezia Tosato, Flora Weissgerber, Laurent Wendling, Sylvain Lobry

机构 * LIPADE, Université Paris Cité(巴黎Cité大学LIPADE研究所) GENCI-IDRIS French Aerospace Lab, ONERA(法国航空航天实验室ONERA)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted at IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 13 pages, 6 figures

Journal ref 10.1109/JSTARS.2025.3596678

详情

展开后加载摘要…

URL PDF HTML 收藏