arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-17 至 2025-11-17 共收录 8 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 8 篇

2511.11438 2025-11-17 cs.CV 79%

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

Mingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li, Yue Qiu, Zekang Du, Mengyang Wu, Pingping Zhang, Kun Li, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments This is the extended version of the paper accepted at AAAI 2026, which includes all technical appendices and additional experimental details

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10705 2025-11-17 cs.AI cs.CL 79%

Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents

Yuan Zhao, Hualei Zhu, Tingyu Jiang, Shen Li, Xiaohang Xu, Hao Henry Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03662 2025-11-17 cs.CV cs.AI cs.RO 73%

Zero-Shot Temporal Interaction Localization for Egocentric Videos

Erhang Zhang, Junyi Ma, Yin-Dong Zheng, Yixuan Zhou, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11502 2025-11-17 cs.CV cs.AI 62%

PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models

Nhat Hoang-Xuan, Minh Vu, My T. Thai, Manish Bhattarai

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) University of Florida(佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11450 2025-11-17 cs.CV cs.LG 62%

VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation

Maximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee, Benjamin Hamm, Constantin Ulrich, Sebastian Regnery, Lukas Bauer, Efthimios Katsigiannopulos, Tobias Norajitra, Klaus Maier-Hein

机构 * German Cancer Research Center, Division of Medical Image Computing(德国癌症研究中心,医学影像计算部) Faculty of Mathematics and Computer Science(数学与计算机科学系) Medical Faculty - Heidelberg University(海德堡大学医学系) Helmholtz Imaging(海德堡影像技术) Department of Radiation Oncology, Heidelberg University Hospital(海德堡大学医院放射肿瘤科) HIDSS4Health, Heidelberg(HIDSS4Health,海德堡) Pattern Analysis and Learning Group, Heidelberg University Hospital(海德堡大学医院模式分析与学习组)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11212 2025-11-17 cs.CV 57%

MAFM^3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI

Mohammad Areeb Qazi, Munachiso S Nwadike, Ibrahim Almakky, Mohammad Yaqub, Numan Saeed

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10923 2025-11-17 cs.CV 57%

Out-of-Distribution Detection with Positive and Negative Prompt Supervision Using Large Language Models

Zhixia He, Chen Zhao, Minglai Shao, Xintao Wu, Xujiang Zhao, Dong Li, Qin Tian, Linlin Yu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10914 2025-11-17 cs.CV 57%

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

Zihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu, Hua Zhang, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Mathematical Sciences, Fudan University(复旦大学数学学院) School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络科学与技术学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏