arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2411.18207 2026-02-27 cs.CV cs.AI 76%

From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects

从开放词汇到开放世界:教会视觉语言模型检测新物体

Zizhao Li, Zhengkang Xiang, Joseph West, Kourosh Khoshelham

机构 * The University of Melbourne Parkville, VIC, Australia(墨尔本大学帕克维尔分校)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV、cs.AI

AI总结 本文提出了一种开放世界框架,使OVD模型能够检测新物体,通过引入OWEL和MSCAL方法提升模型对远超出分布物体的识别能力。

Comments Accepted by BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13314 2026-02-25 cs.CV cs.AI 76%

Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction

Sim2Radar:通过VLM引导的场景重建弥合雷达仿真到现实的差距

Emily Bejerano, Federico Tondolo, Ayaan Qayyum, Xiaofan Yu, Xiaofan Jiang

机构 * Columbia University(哥伦比亚大学) University of California, Merced(加州大学默塞德分校)

专题命中 视觉定位与Grounding :VLM(title);分类 cs.CV、cs.AI

AI总结 Sim2Radar通过VLM引导的场景重建,利用合成数据提升雷达感知性能,实现3D AP提升3.7个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18487 2026-02-17 cs.RO cs.CV cs.LG 76%

Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning

在视觉表示中 grounding 身体意识以实现高效的策略学习

Junlin Wang, Zhiyun Lin

机构 * SUSTech Shenzhen(深圳科技大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

AI总结 本文提出ICon方法,通过对比学习提升机器人操作中策略学习的效率和跨机器人迁移能力。

Comments A preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21296 2026-01-30 cs.LG cs.AI 76%

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

基于信息性和效用的蒸馏数据集 grounding

Shaobo Wang, Yantai Yang, Guo Chen, Peiru Li, Kaixin Li, Yufa Zhou, Zhaorun Chen, Linfeng Zhang

机构 * EPIC Lab, SJTU(SJTU实验室) Shanghai Jiao Tong University(上海交通大学) National University of Singapore(新加坡国立大学) Duke University(杜克大学) The University of Chicago(芝加哥大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 本文提出InfoUtil框架,通过博弈论和梯度范数优化,提升数据集蒸馏的信息性和效用,实验显示在ImageNet-1K上性能提升6.1%。

Comments Accepted by ICLR 2026, 20 pages, 9 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19366 2025-12-11 cs.LG cs.AI 76%

Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs

未接地的接地:一种基于谱-图框架的多模态大语言模型幻觉量化方法

Supratik Sarkar, Swagatam Das

机构 * Morgan Stanley(摩根士丹利) Indian Statistical Institute (Kolkata)(印度统计研究所(加尔各答))

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于谱-图框架的多模态大语言模型幻觉量化方法,通过谱分解和RKHS本征模式,提供可解释的语义失真度量。

Comments 49 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06040 2025-10-08 cs.CV cs.AI 76%

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Chao Wang, Yuqi Pan, Tianhao Hou, Xiaojuan Wang, Yutong Gao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Minzu University of China(民族大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 76%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12496 2025-07-18 cs.RO cs.AI cs.LG 76%

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making

Yucen Wang, Rui Yu, Shenghua Wan, Le Gan, De-Chuan Zhan

机构 * School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Accepted by Forty-Second International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16826 2025-06-23 cs.CV cs.AI cs.RO 76%

AnyTraverse: An off-road traversability framework with VLM and human operator in the loop

Sattwik Sahu, Agamdeep Singh, Karthik Nambiar, Srikanth Saripalli, P. B. Sujit

机构 * Department of Electrical Engineering and Computer Science, Indian Institute of Science Education and Research Bhopal(电子工程与计算机科学系,印度科学教育与研究学院布达佩斯分校) Department of Mechanical Engineering, Texas A&M University(机械工程系,德克萨斯A&M大学)

专题命中 视觉定位与Grounding :VLM(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00412 2025-06-23 cs.LG cs.AI cs.CL 76%

Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation

Bohan Lyu, Yadi Cao, Duncan Watson-Parris, Leon Bergen, Taylor Berg-Kirkpatrick, Rose Yu

机构 * Tsinghua University(清华大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 37 pages, 16 figures

Journal ref In Proceedings of the Forty-second International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20252 2025-05-21 cs.CV cs.AI 76%

LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions

Yejin Kwon, Daeun Moon, Youngje Oh, Hyunsoo Yoon

机构 * Department of Industrial Engineering, Yonsei University(工业工程系,延世大学)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV、cs.AI

Comments Accepted Industry Track at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16047 2025-04-23 cs.CV cs.AI 76%

Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis

Frank Li, Hari Trivedi, Bardia Khosravi, Theo Dapamede, Mohammadreza Chavoshi, Abdulhameed Dere, Rohan Satya Isaac, Aawez Mansuri, Janice Newsome, Saptarshi Purkayastha, Judy Gichoya

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16488 2025-03-26 cs.HC cs.CV cs.LG 76%

VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection

Kunal Chavan, Keertan Balaji, Spoorti Barigidad, Samba Raju Chiluveru

机构 * IIT Dharwad(达尔瓦德印度理工学院) S.R.M Institute of Technology(SRM理工学院)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02917 2025-03-06 eess.IV cs.AI cs.CV 76%

Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models

Deval Mehta, Yiwen Jiang, Catherine L Jan, Mingguang He, Kshitij Jadhav, Zongyuan Ge

机构 * Monash University(莫纳什大学) Centre for Eye Research Australia(澳大利亚眼科研究中心) Royal Victorian Eye and Ear Hospital(皇家维多利亚眼耳医院) The University of Melbourne(墨尔本大学) The Hong Kong Polytechnic University(香港理工大学) Centre for Eye and Vision Research(眼与视觉研究中心) Indian Institute of Technology Bombay(印度理工学院孟买分校) Airdoc-Monash Research Lab(Airdoc-莫纳什研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted to Information Processing in Medical Imaging (IPMI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11168 2025-02-18 cs.CV cs.AI 76%

Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

Xin Gu, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang, Yuewei Lin, Heng Fan, Libo Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) La Trobe University(拉筹伯大学) University of North Texas(北得克萨斯大学) Brookhaven National Laboratory(布鲁克海文国家实验室)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03722 2025-01-08 cs.CV cs.AI 76%

Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein

Xiaotong Guo, Deqian Yang, Dan Wang, Haochen Zhao, Yuan Li, Zhilin Sui, Tao Zhou, Lijun Zhang, Yanda Meng

机构 * National Clinical Research Center for Cancer/Cancer Hospital Shenzhen Hospital(国家癌症临床研究中心/肿瘤医院深圳医院) Key Laboratory of System Software (Chinese Academy of Sciences)(系统软件重点实验室(中国科学院)) State Key Laboratory of Computer Science(计算机科学国家重点实验室) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Guangzhou Jiayi Software Technology Co., Ltd.(广州嘉益软件科技有限公司) R&D Center, Guangxi Huayi Artificial Intelligence Medical Technology Co., Ltd(广西华奕人工智能医疗科技有限公司研发中心) Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Department of Cardiovascular & Metabolic Medicine, University of Liverpool(利物浦大学心血管与代谢医学系)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments 8 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23214 2024-11-01 cs.LG cs.AI 76%

Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval

Sheryl Hsu, Omar Khattab, Chelsea Finn, Archit Sharma

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03959 2024-10-08 cs.CL cs.AI cs.CV cs.GR 76%

Grounding Language in Multi-Perspective Referential Communication

Zineng Tang, Lingjun Mao, Alane Suhr

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted to EMNLP2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18868 2024-09-30 cs.CL cs.AI cs.LG 76%

Individuation in Neural Models with and without Visual Grounding

Alexey Tikhonov, Lisa Bylinina, Ivan P. Yamshchikov

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01326 2024-09-04 cs.RO cs.AI cs.LG 76%

Grounding Language Models in Autonomous Loco-manipulation Tasks

Jin Wang, Nikos Tsagarakis

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Summit to ICRA@40. arXiv admin note: substantial text overlap with arXiv:2406.14655

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02244 2024-08-06 cs.CV cs.AI 76%

Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets

Lucas Choi, Ross Greer

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03310 2024-07-19 cs.AI cs.CV 76%

V-IRL: Grounding Virtual Intelligence in Real Life

Jihan Yang, Runyu Ding, Ellis Brown, Xiaojuan Qi, Saining Xie

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Project page: https://virl-platform.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01256 2024-06-04 cs.CV cs.AI 76%

Augmented Commonsense Knowledge for Remote Object Grounding

Bahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu, Shirui Pan, Javen Qinfeng Shi

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19336 2024-03-29 cs.CV cs.AI 76%

IVLMap: Instance-Aware Visual Language Grounding for Consumer Robot Navigation

Jiacui Huang, Hongtao Zhang, Mingbo Zhao, Zhou Wu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03936 2023-12-08 cs.CV cs.CY cs.LG cs.SI 76%

The Potential of Vision-Language Models for Content Moderation of Children's Videos

Syed Hammad Ahmed, Shengnan Hu, Gita Sukthankar

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.LG

Comments 5 pages, 1 figure. Accepted at IEEE ICMLA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01656 2023-12-06 cs.IR cs.AI cs.CV cs.HC 76%

The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model

Yilin Ye, Qian Zhu, Shishi Xiao, Kang Zhang, Wei Zeng

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted by The 2024 ACM SIGCHI Conference on Computer-Supported Cooperative Work & Social Computing (CSCW) (Proc. CSCW 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00729 2023-11-08 cs.CV cs.AI 76%

ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection

Thinh Phan, Khoa Vo, Duy Le, Gianfranco Doretto, Donald Adjeroh, Ngan Le

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05186 2023-08-22 cs.CV cs.AI 76%

Suspected Object Matters: Rethinking Model's Prediction for One-stage Visual Grounding

Yang Jiao, Zequn Jie, Jingjing Chen, Lin Ma, Yu-Gang Jiang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted to ACM MM 23

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10438 2023-05-19 cs.CL cs.AI cs.CV cs.MM 76%

IMAGINATOR: Pre-Trained Image+Text Joint Embeddings using Word-Level Grounding of Images

Varuna Krishna, S Suryavardan, Shreyash Mishra, Sathyanarayanan Ramamoorthy, Parth Patwa, Megha Chakraborty, Aman Chadha, Amitava Das, Amit Sheth

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07528 2023-05-15 cs.CV cs.AI 76%

WEDGE: A multi-weather autonomous driving dataset built from generative vision-language models

Aboli Marathe, Deva Ramanan, Rahee Walambe, Ketan Kotecha

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted in Vision Datasets Understanding at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏