arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26403 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2603.02329 2026-03-04 cs.CV 89%

HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

HAMMER: 通过跨模态整合利用大语言模型进行意图驱动的3D affordance grounding

Lei Yao, Yong Chen, Yuejiao Su, Yi Wang, Moyun Liu, Lap-Pui Chau

机构 * The Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 HAMMER通过跨模态整合多模态大语言模型,实现意图驱动的3D affordance grounding,提升3D表示的准确性和鲁棒性。

Comments Accepted by CVPR 2026. Project Page: https://rayyoh.github.io/Hammer

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20794 2026-02-25 cs.CV 89%

VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

VGGDrive: 通过跨视角几何 grounding 为自动驾驶赋能 Vision-Language 模型

Jie Wang, Guang Li, Zhijian Huang, Chenxu Dang, Hangjun Ye, Yahong Han, Long Chen

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) Xiaomi EV(小米汽车)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV

AI总结 VGGDrive通过引入跨视角几何 grounding 机制,提升Vision-Language模型在自动驾驶任务中的性能表现。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11904 2025-05-13 cs.CV 89%

GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

Yue Zhou, Mengcheng Lan, Xiang Li, Litong Feng, Yiping Ke, Xue Jiang, Qingyun Li, Xue Yang, Wayne Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Reading(阅读大学) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学) SenseTime Research(商汤科技研究院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13983 2025-04-14 cs.CV 89%

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li, Jiannan Ge, Hongtao Xie, Yongdong Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19325 2025-03-14 cs.CV 89%

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, Salman Khan

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(title,abstract);LLaVA(abstract);分类 cs.CV

Comments This updated version includes revisions and additional analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10419 2025-03-14 cs.RO cs.AI 89%

HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models

Vineet Bhat, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12694 2025-02-19 cs.CV cs.CL 89%

VividMed: Vision Language Model with Versatile Visual Grounding for Medicine

Lingxiao Luo, Bingda Tang, Xuanzhong Chen, Rong Han, Ting Chen

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10840 2024-12-17 cs.CV 89%

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning

Hai-Ming Xu, Qi Chen, Lei Wang, Lingqiao Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14901 2024-11-25 cs.CV cs.CL 89%

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu, Thomas Seidl, Gedas Bertasius

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14492 2024-06-21 cs.CV cs.CL 89%

Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?

Gregor Geigle, Radu Timofte, Goran Glavaš

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12537 2023-08-25 cs.RO cs.CV 89%

HuBo-VLM: Unified Vision-Language Model designed for HUman roBOt interaction tasks

Zichao Dong, Weikun Zhang, Xufeng Huang, Hang Ji, Xin Zhan, Junbo Chen

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(title);vision language model(abstract);grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14824 2023-07-14 cs.CL cs.CV 89%

Kosmos-2: Grounding Multimodal Large Language Models to the World

Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Furu Wei

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24934 2026-08-27 cs.CV cs.AI 新提交 89%

Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

融合感知视觉专家与多模态大语言模型的可解释植物病害诊断:从基准图像到真实世界机器人现场验证

Ranjan Sapkota, Konstantinos I. Roumeliotis, Pengyao Xie, Nikolaos D. Tselikas, Lirong Xiang, Manoj Karkee

机构 * Cornell University(康奈尔大学) University of the Peloponnese(伯罗奔尼撒大学) Agricultural University of Athens(雅典农业大学)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

AI总结 该研究提出H²MAF框架,结合EfficientNet-B3、ConvNeXt-Tiny与Gemma、Qwen等MLLM,在PlantDoc及Cornell机器人田间数据集上实现高准确率可解释植物病害诊断,验证了MLLM仲裁的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15517 2026-07-20 cs.CV cs.AI 新提交 89%

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

SLAPBench:用于四指SLAP指纹验证的多模态大语言模型基准测试

Bibesh Pyakurel, M. G. Sarwar Murshed

机构 * University of Wisconsin–Green Bay(威斯康星大学格林湾分校)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

AI总结 研究四指SLAP指纹验证,介绍SLAPBench基准,评估多个MLLM在不同提示下的表现,发现提示控制崩溃,模型能力控制歧视,建立了特定于SLAP的MLLM基线,揭示了模型在指纹验证中的能力差距和公平性问题。

Comments 19 pages, 6 figures, 2 tables. Includes appendix with supporting figures and per-subgroup fairness detail. Code and data: https://github.com/bibeshpyakurel/SLAPBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03553 2026-06-04 cs.CV cs.AI 89%

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching

直播中的动态内容审核:结合监督分类与MLLM增强的相似度匹配

Wei Chee Yew, Hailun Xu, Sanjay Saha, Xiaotian Fan, Hiok Hian Ong, David Yuchen Wang, Kanchan Sarkar, Zhenheng Yang, Danhui Guan

机构 * TikTok Singapore Singapore(TikTok新加坡) TikTok San Jose United States(TikTok旧金山美国) TikTok Shanghai China(TikTok上海中国)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 提出一种混合审核框架,结合监督分类和基于参考的相似度匹配,利用多模态大语言模型提升准确性,在保持轻量推理的同时实现大规模直播内容审核。

Comments To be published at KDD 2026 (ADS track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23797 2026-05-25 cs.LG cs.CV 89%

Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models

去偏负挖掘提升基于预训练视觉语言模型的分布外检测

Bo Peng, Jie Lu, Guangquan Zhang, Zhen Fang

机构 * University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);分类 cs.CV、cs.LG

AI总结 针对分布外检测中负标签的假阴性问题,提出通过间接近似负标签分布来校正采样偏差的理论框架,并转化为基于ID标签和未标注语料数据的蒙特卡洛采样方法,在多种OOD检测设置中达到新最优。

Comments KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09879 2026-01-16 cs.CV cs.AI 89%

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割

Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong

机构 * Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系) Department of Radiology, University of Florida(佛罗里达大学放射学系) Research Computing, University of Florida(佛罗里达大学研究计算中心) Department of Medicine, University of Florida(佛罗里达大学医学系) Department of Radiology, UC San Francisco(旧金山大学放射学系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);visual reasoning(abstract);visual question answering(abstract)

AI总结 MedVL-SAM2是一种统一的3D医学多模态模型,通过联合训练实现报告生成、VQA和多任务分割的高性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14732 2026-08-17 cs.LG cs.AI cs.CV eess.IV 版本更新 89%

INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT

INFORM-CT:整合LLM和VLM用于腹部CT的偶发发现管理

Idan Tankel, Nir Mazor, Rafi Brada, Christina LeBedis, Guy ben-Yosef

机构 * GE Healthcare Technology and Innovation Center(GE医疗技术与创新中心) Boston Medical Center(波士顿医疗中心)

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出基于LLM和VLM的计划-执行框架,用于提高腹部CT偶发发现的检测、分类和报告效率与精度,通过自动化流程提升临床应用效果。

Comments Spotlight presentation at the 9th International Conference on Medical Imaging with Deep Learning (MIDL) 2026 Additional code and implementation details available at https://idan-tankel.github.io/InformCT_ProjectPage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19130 2026-05-20 cs.LG cs.AI cs.CL cs.CV 89%

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

EgoBabyVLM:基于自然主义第一人称视频数据的跨模态学习基准测试

Dongyan Lin, Phillip Rust, Angel Villar Corrales, Alvin W. M. Tan, Mahi Luthra, Charles-Éric Saint-James, Rashel Moritz, Sheila Krogh-Jespersen, Vanessa Stark, Surya Parimi, Jiayi Shen, Youssef Benchekroun, Yosuke Higuchi, Martin Gleize, Tom Fizycki, Nicolas Hamilakis, Manel Khentout, Sho Tsuji, Balázs Kégl, Juan Pino, Michael C. Frank, Emmanuel Dupoux

机构 * Meta Superintelligence Labs(Meta超智能实验室) Stanford University(斯坦福大学) Meta Reality Labs(Meta现实实验室) The University of Tokyo(东京大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究探讨了儿童如何从有限的视觉-语言输入中获得语言 grounding 的鲁棒性,提出了 EgoBabyVLM 挑战,推动模型在自然主义数据中实现 grounded language learning。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18988 2026-03-20 cs.RO 89%

MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction

MERGE:面向人类机器人交互中多主体事件推理与 grounding 的引导视觉-语言模型

Joerg Deigmoeller, Nakul Agarwal, Stephan Hasler, Daniel Tanneberg, Anna Belardinelli, Reza Ghoddoosian, Chao Wang, Felix Ocker, Fan Zhang, Behzad Dariush, Michael Gienger

机构 * Honda Research Institute Europe(本田欧洲研究院) Honda Research Institute USA(本田美国研究院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract)

AI总结 MERGE 通过引导视觉语言模型实现动态人类机器人交互中多主体事件的推理与 grounding,提升 situational awareness 与效率,基于 GROUND 数据集验证其性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15333 2025-11-20 cs.RO 89%

C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models

Nayoung Oh, Dohyun Kim, Junhyeong Bang, Rohan Paul, Daehyung Park

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract)

Comments 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04974 2025-09-25 cs.CV cs.AI cs.CL cs.LG 89%

Towards Visual Text Grounding of Multimodal Large Language Model

Ming Li, Ruiyi Zhang, Jian Chen, Chenguang Wang, Jiuxiang Gu, Yufan Zhou, Franck Dernoncourt, Wanrong Zhu, Tianyi Zhou, Tong Sun

机构 * Adobe Research(Adobe研究院) University of Maryland(马里兰大学) University at Buffalo(布法罗大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05861 2024-04-03 cs.CL cs.AI cs.CV cs.LG 89%

Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models

Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG

Comments ICLR 2024 camera-ready (23 pages), Code: https://github.com/archiki/RepARe

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02325 2024-03-05 cs.CV cs.AI cs.CL cs.LG 89%

Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training

David Wan, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project website: https://contrastive-region-guidance.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21180 2026-08-24 eess.IV cs.CV 新提交 89%

Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

面向基于视觉语言模型的LGE-MR图像临床质量与可用性评估以用于心脏消融规划

Bipasha Kundu, Abhishek Chaturvedi, Axel W. E. Wismueller, Richard Simon, Cristian A. Linte

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision language model(title,abstract);分类 cs.CV

AI总结 本研究提出两阶段VLM框架用于左心房LGE-MRI的临床导向图像质量评估,在60个图像切片-文本对数据集上,InternVL2的标准级准确率最高,DeepSeek实现完美临床可用性一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18309 2026-08-20 cs.CV 新提交 89%

XRF-to-Optical Field-of-View Localization with Vision Language Models

基于视觉语言模型的X射线荧光(XRF)与光学显微镜视场(FOV)定位

Xiangyu Yin, Tatjana Paunesku, Letonia Copeland-Hardin, Martina Ralle, Zichao Wendy Di, Si Chen, Gayle E. Woloschak, Barry Lai, Mathew J. Cherukara, Stefan Vogt

机构 * Northwestern University(西北大学) University of Chicago(芝加哥大学) Oregon Health and Science University(俄勒冈健康与科学大学) Argonne National Laboratory(阿贡国家实验室)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision language model(title,abstract);分类 cs.CV

AI总结 本文针对跨模态显微图像的视场定位难题,提出结合视觉语言模型(VLM)的候选生成-验证工作流,在低对应度的相邻切片成像数据中实现了有效定位,为关联XRF与光学显微测量提供了支撑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09302 2026-08-12 cs.CV 版本更新 89%

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

用于宫腔镜手术场景分割的自举式视觉-语言模型

Jun Huang, Meiyi Chen, Zijie Yue, Yuhang Xiao, Fang Li, Hanli Wang, Xiaowen Tong, Yi Guo, Miaojing Shi

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);分类 cs.CV

AI总结 本研究提出首个基于VLM的宫腔镜手术场景分割方法VLM-hyster,通过类别特定文本提示与掩码蒸馏分支提升性能,在自行构建的4020张图像数据集上表现优于现有模型,获多中心验证,具临床应用潜力。

Comments Accept by Biomedical Signal Processing and Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09691 2026-08-11 cs.CV 新提交 89%

Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

扩散目标,保留其标签:通过VLM构建的3D植被场景从少量未标记照片中整理检测器训练数据

Mario Malizia, Marnix Enting, Rob Haelterman, Ken Hasselmann

机构 * Royal Military Academy(皇家军事学院) KU Leuven(鲁汶大学) Flanders Make(佛兰德制造研究院)

专题命中 视觉定位与Grounding :VLM(title,title_cn);分类 cs.CV

AI总结 本研究针对植被中小物体标记图像稀缺导致检测器跨站点泛化差的问题,通过VLM构建3D植被场景合成标记训练图像,实现无监督站点适应,在排雷基准上性能优于传统跨站点标签复用。

Comments Accepted at the Curated Data for Efficient Learning (CDEL) Workshop @ ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01258 2026-08-04 cs.CV 新提交 89%

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

多模态大语言模型生成图像检测的基准数据集:GPT Image2与Nano Banana2

Zirui Zhang, Yinbo Yu, Donghai Guan, Chunwei Tian, Daoqiang Zhang, Qi Zhu

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :MLLM(title,summary_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文构建了含三种生成协议的MLLM生成图像检测基准数据集,评估现有检测器性能并提出SAP-DSP基线框架,验证了其检测稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05161 2026-06-23 cs.CV 版本更新 89%

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

基于VLM的医学图像分布外检测的Wasserstein对齐定位

Bernhard Kainz, Johanna P Mueller, Matthew Baugh, Cosmin Bercea

机构 * Department of Computing, Imperial College London, UK(伦敦帝国理工学院计算机系) Technical University Munich, DE(慕尼黑技术大学) Munich Center for Machine Learning (MCML), DE(慕尼黑机器学习中心(MCML))

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);visual reasoning(abstract);分类 cs.CV

AI总结 提出WALDO框架,利用最优传输理论通过熵加权切片Wasserstein距离、Goldilocks区域采样和自一致性聚合实现零样本异常定位,在NOVA脑MRI基准上mAP@30达43.5%,相对提升19%。

Comments submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏