arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-14 至 2025-11-14 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 5 篇

2511.10059 2025-11-14 cs.CV 81%

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, Yu Zhou

专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10591 2025-11-14 cs.CL cs.AI 74%

Mined Prompting and Metadata-Guided Generation for Wound Care Visual Question Answering

Bavana Durgapraveen, Sornaraj Sivasankaran, Abhinand Balachandran, Sriram Rajkumar

机构 * EXL Health AI Lab at MEDIQA-WV 2025(EXL健康AI实验室)

专题命中 视觉问答 :visual question answering(title);分类 cs.AI

Comments 2 figures, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09651 2025-11-14 cs.IR cs.AI cs.CL cs.LG eess.SP 62%

Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations

Zakaria El Kassimi, Fares Fourati, Mohamed-Slim Alouini

机构 * KAUST(卡斯泰尔大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 7 figures, AI4NextG @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09868 2025-11-14 cs.CV 57%

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen, Hui Zhang

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 57%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏