arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2511.21272 2026-01-12 cs.CV 85%

Co-Training Vision Language Models for Remote Sensing Multi-task Learning

联合训练遥感多任务学习的视觉语言模型

Qingyun Li, Shuran Ma, Junwei Luo, Yi Yu, Yue Zhou, Fengxiang Wang, Xudong Lu, Xiaoxing Wang, Xin He, Yushi Chen, Xue Yang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Shanghai Jiao Tong University(上海交通大学) Xidian University(西安电子科技大学) East China Normal University(东华大学) Wuhan University(武汉大学) Southeast University(东南大学) National University of Defense Technology(国防科技大学) Chinese University of Hong Kong(香港中文大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 RSCoVLM通过联合训练视觉语言模型,提升遥感多任务学习的性能和灵活性。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21402 2025-12-29 cs.CV 85%

Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation

理解病毒性:一种基于Rubric的视觉-语言模型框架用于短形式教育娱乐内容评估

Arnav Gupta, Gurekas Singh Sahney, Hardik Rathi, Abhishek Chandwani, Ishaan Gupta, Pratik Narang, Dhruv Kumar

机构 * Birla Institute of Technology and Science, Pilani(比拉理工学院和科学学院,比里尼)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出基于Rubric的视觉-语言模型框架,用于评估短形式教育娱乐视频的病毒性,通过提取音频视觉特征并训练回归评估器,实现可解释且可扩展的参与度预测。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18947 2025-12-23 cs.CV 85%

OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model

OpenHOI: 基于多模态大语言模型的开放世界手-物体交互合成

Zhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni, Qi Ye, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

AI总结 OpenHOI通过多模态大语言模型实现开放世界手-物体交互合成,能生成长周期操控序列并处理复杂语言指令。

Comments Accepted by NeurIPS 2025 as Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11099 2025-12-15 cs.CV 85%

VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction

VGent: 通过模块化设计实现视觉 grounding 的解耦推理与预测

Weitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei, Yan Yan, Kangning Liu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Adobe(Adobe公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 VGent 通过模块化设计实现视觉 grounding 的解耦推理与预测,利用冻结 MLLM 和解码器交叉注意力机制,提升多目标识别性能。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09215 2025-12-11 cs.CV 85%

View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs

视图-图:通过场景图上的视觉-语言推理实现零样本3D视觉定位

Yuanyuan Liu, Haiyang Mei, Dongyang Zhan, Jiayue Zhao, Dongsheng Zhou, Bo Dong, Xin Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出视图-图方法,通过场景图上的视觉-语言推理实现零样本3D视觉定位,有效解决空间语义关系处理难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21375 2025-11-27 cs.CV 85%

Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning

通过边界框思考:通过强化微调提升空间时间视频定位

Xin Gu, Haoji Zhang, Qihang Fan, Jingxuan Niu, Zhipeng Zhang, Libo Zhang, Guang Chen, Fan Chen, Longyin Wen, Sijie Zhu

机构 * ByteDance Intelligent Creation(字节跳动智能创作) Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai Jiao Tong University(上海交通大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 STVG-o1通过强化微调提升空间时间视频定位性能,实现无需架构修改的先进结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19516 2025-11-27 cs.CV 85%

Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning

连接点:通过代理推理实现无需训练的视觉接地

Liqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai, Yixiong Zou, Yonghong Tian

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 GroundingAgent通过代理推理实现无需微调的视觉接地,达到65.1%的零样本准确率,并在选择阶段实现90%的准确率,展示了LLM推理能力的重要性。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10117 2025-11-11 cs.CV 85%

AGO: Adaptive Grounding for Open World 3D Occupancy Prediction

Peizheng Li, Shuxiao Ding, You Zhou, Qingwen Zhang, Onat Inak, Larissa Triess, Niklas Hanselmann, Marius Cordts, Andreas Zell

机构 * Mercedes-Benz AG(梅赛德斯-奔驰集团) University of Tübingen(图宾根大学) Tübingen AI Center(图宾根人工智能中心) University of Bonn(波恩大学) RPL KTH Royal Institute of Technology(皇家理工学院) TU Berlin(柏林技术大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23449 2025-10-28 cs.MM cs.CV cs.IR 85%

CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection

Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou

机构 * School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) National University of Singapore(新加坡国立大学) Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)

专题命中 视觉定位与Grounding :MLLM(title,abstract);visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16538 2025-10-21 cs.CV cs.RO 85%

Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

Bastian Pätzold, Jan Nogga, Sven Behnke

机构 * Autonomous Intelligent Systems, University of Bonn(博恩大学自主智能系统中心) Lamarr Institute for Machine Learning and AI(拉马尔人工智能与机器学习研究所) Center for Robotics, University of Bonn(博恩大学机器人中心)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments IEEE Robotics and Automation Letters (RA-L), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16290 2025-10-21 cs.CV cs.CL 85%

Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models

Yue Zheng, Xiufang Shi, Jiming Chen, Yuanchao Shu

机构 * Zhejiang University of Technology(浙江工业大学) Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14672 2025-10-17 cs.CV 85%

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

Jinglei Zhang, Yuanfan Guo, Rolandos Alexandros Potamias, Jiankang Deng, Hang Xu, Chao Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Noah’s Ark Lab(诺亚实验室) Imperial College London(伦敦帝国理工学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07975 2025-10-10 cs.RO cs.AI 85%

Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation

Mingyang Sun, Jiude Wei, Qichen He, Donglin Wang, Cewu Lu, Jianhua Sun

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02750 2025-10-06 cs.CV 85%

Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models

Lihua Zhou, Mao Ye, Shuaifeng Li, Nianxin Li, Jinlin Wu, Xiatian Zhu, Lei Deng, Hongbin Liu, Jiebo Luo, Zhen Lei

机构 * Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China(人工智能与机器人研究中心,香港科学与创新研究院,中国科学院,香港,中国) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科学与技术大学计算机科学与工程学院) Surrey Institute for People-Centred Artificial Intelligence, CVSSP, University of Surrey(以人为中心的人工智能研究院,CVSSP, Surrey大学) School of Electronics and Information Engineering, Shenzhen University(电子与信息工程学院,深圳大学) University of Rochester(罗切斯特大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24133 2025-09-30 cs.CV cs.CL 85%

Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding

Zhecheng Li, Guoxian Song, Yiwei Wang, Zhen Xiong, Junsong Yuan, Yujun Cai

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) University of California, Merced(加州大学默塞德分校) University of Southern California(南加州大学) University at Buffalo(布法罗大学) The University of Queensland(昆士兰大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06798 2025-08-07 cs.CV 85%

Uncertainty-aware Medical Diagnostic Phrase Identification and Grounding

Ke Zou, Yang Bai, Bo Liu, Yidi Chen, Zhihao Chen, Yang Zhou, Xuedong Yuan, Meng Wang, Xiaojing Shen, Xiaochun Cao, Yih Chung Tham, Huazhu Fu

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Department of Radiology, West China Hospital, Sichuan University(四川大学华西医院放射科) College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院) Department of Mathematics, Sichuan University(四川大学数学系) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院) Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore and the Singapore Eye Research Institute, Singapore National Eye Centre(新加坡国立大学 Yong Loo Lin 医学院眼科系及新加坡眼科研所、新加坡国家眼科中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02890 2025-08-06 cs.CV cs.CL 85%

VisuCraft: Enhancing Large Vision-Language Models for Complex Visual-Guided Creative Content Generation via Structured Information Extraction

Rongxin Jiang, Robert Long, Chenghao Gu, Mingrui Yan

机构 * Heilongjiang University of Science and Technology(黑龙江科技大学) University of Padua(帕多瓦大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01008 2025-08-05 cs.CV 85%

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation

Cihang Peng, Qiming Hou, Zhong Ren, Kun Zhou

机构 * State Key Lab of CAD&CG(计算机辅助设计与图形学国家重点实验室)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21316 2025-07-17 cs.CV 85%

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents

Badri Vishal Kasuba, Parag Chaudhuri, Ganesh Ramakrishnan

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10596 2025-07-16 cs.CV 85%

GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding

Rui Hu, Lianghui Zhu, Yuxuan Zhang, Tianheng Cheng, Lei Liu, Heng Liu, Longjin Ran, Xiaoxin Chen, Wenyu Liu, Xinggang Wang

机构 * School of EIC, Huazhong University of Science & Technology(华中科技大学电子信息学院) vivo AI Lab(vivo人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments To appear at ICCV 2025. Code: https://github.com/hustvl/GroundingSuite

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19498 2025-06-25 cs.RO cs.AI 85%

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

Yiteng Chen, Wenbo Li, Shiyi Wang, Huiping Zhuang, Qingyao Wu

机构 * School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学) Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI

Comments submitted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17901 2025-06-24 cs.CV 85%

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

Yixuan Wu, Yang Zhang, Jian Wu, Philip Torr, Jindong Gu

机构 * University of Oxford(牛津大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20104 2025-06-16 cs.CV 85%

New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration

Xuzheng Yang, Junzhuo Liu, Peng Wang, Guoqing Wang, Yang Yang, Heng Tao Shen

机构 * University of Electronic Science and Technology of China(电子科技大学) Tongji University(同济大学)

专题命中 视觉定位与Grounding :MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01977 2025-06-10 cs.CV 85%

AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs

Hongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen, Qing Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院模式识别新技术实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室) Hong Kong Institute of Science & Innovation, CASIA(香港科学与创新研究院) The Hong Kong Polytechnic University(香港理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22222 2025-05-29 cs.CV cs.CL 85%

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

Yunsoo Kim, Jinge Wu, Su-Hwan Kim, Pardeep Vasudev, Jiashu Shen, Honghan Wu

机构 * UCL(伦敦大学学院) Technical University of Munich(慕尼黑技术大学) University of Oxford(牛津大学) University of Glasgow(格拉斯哥大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19552 2025-05-23 cs.CV 85%

GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Hosam Elgendy, Ahmed Sharshar, Ahmed Aboeitta, Yasser Ashraf, Mohsen Guizani

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎德·本·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV

Comments 14 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14231 2025-05-21 cs.CV 85%

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Sule Bai, Mingxing Li, Yong Liu, Jing Tang, Haoji Zhang, Lei Sun, Xiangxiang Chu, Yansong Tang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) AMAP, Alibaba Group(阿里云研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11852 2025-05-20 cs.CV 85%

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding

Jingkun Yue, Siqi Zhang, Zinan Jia, Huihuan Xu, Zongbo Han, Xiaohong Liu, Guangyu Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tianjin University(天津大学) South China Hospital, Medical School, Shenzhen University(深圳大学医学院南方医院)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02278 2025-05-06 cs.CV 85%

Compositional Image-Text Matching and Retrieval by Grounding Entities

Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted at CVPR-W

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01578 2025-05-06 cs.CV 85%

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

Gabriel Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, Andrew D. Wilson

机构 * Microsoft Research(微软研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏