arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-07 至 2025-11-07 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 5 篇

2511.04384 2025-11-07 cs.CV cs.LG 74%

Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA

Itbaan Safwan, Muhammad Annas Shaikh, Muhammad Haaris, Ramail Khan, Muhammad Atif Tahir

机构 * School of Mathematics and Computer Science, Institute of Business Administration (IBA), Karachi, Pakistan(数学与计算机科学学院,商学院(IBA),巴基斯坦卡里奇)

专题命中 视觉问答 :visual question answering(abstract,comments);grounding(abstract);分类 cs.CV、cs.LG

Comments This is a working paper submitted for Medico 2025: Visual Question Answering (with multimodal explanations) for Gastrointestinal Imaging at MediaEval 2025. 5 pages, 3 figures and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10528 2025-11-07 cs.CV cs.AI 73%

Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset

Ziye Deng, Ruihan He, Jiaxiang Liu, Yuan Wang, Zijie Meng, Songtao Jiang, Yong Xie, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Guangdong Institute of Intelligence Science(广东智能科学研究院)

专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21814 2025-11-07 cs.CV cs.AI 62%

Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding

Zhuoming Li, Aitong Liu, Mengxi Jia, Yubi Lu, Tengxiang Zhang, Changzhi Sun, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究院(TeleAI))

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments IMWUT2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04502 2025-11-07 cs.CL cs.AI 57%

RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG

Joshua Gao, Quoc Huy Pham, Subin Varghese, Silwal Saurav, Vedhus Hoskere

机构 * University of Houston(德克萨斯大学休斯敦分校)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03325 2025-11-07 cs.CV 57%

SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding

Mauro Orazio Drago, Luca Carlini, Pelinsu Celebi Balyemez, Dennis Pierantozzi, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque

机构 * Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB)(电子、信息与生物工程系) Politecnico di Milano(米兰理工大学) IRCCS Humanitas Research Hospital(IRCCS人类itas研究医院) UCL Hawkes Institute and Department of Computer Science(UCL Hawkes研究所和计算机科学系) University College London(伦敦大学学院) University of Manchester(曼彻斯特大学)

专题命中 视觉问答 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏