arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-07-31 至 2025-07-31 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2503.07631 2025-07-31 cs.LG cs.CL 83%

OWLViz: An Open-World Benchmark for Visual Question Answering

Thuy Nguyen, Dang Nguyen, Hoang Nguyen, Thuan Luong, Long Hoang Dang, Viet Dac Lai

机构 * Posts and Telecommunications Institute of Technology, Viet Nam(越南电信技术研究所) Adobe Research, USA(Adobe研究) University of Maryland, USA(美国马里兰大学)

专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.LG

Comments 8 pages + appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22346 2025-07-31 cs.CV 77%

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

Pei Deng, Wenqian Zhou, Hanlin Wu

机构 * School of Information Science and Technology, Beijing Foreign Studies University(信息科学与技术学院,北京外国语大学)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

Comments 12 pages, 5 figures. Submitted to IEEE Transactions on Geoscience and Remote Sensing (TGRS). Code and dataset are available at https://github.com/hanlinwu/DeltaVLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04585 2025-07-31 cs.CL 50%

MuSciClaims: Multimodal Scientific Claim Verification

Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan, Apoorva Kashi, Mahnaz Koupaee, Niranjan Balasubramanian

机构 * Stony Brook University(石溪大学)

专题命中 视觉问答 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏