arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2411.04351 2025-08-01 cs.CV 79%

LidaRefer: Context-aware Outdoor 3D Visual Grounding for Autonomous Driving

Yeong-Seung Baek, Heung-Seon Oh

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19947 2025-07-31 cs.RO cs.CL cs.IT cs.LG cs.SY eess.SY math.IT 79%

Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations

Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich, Rina Tse

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments Accepted to the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC); Supplementary video: https://cu-asl.github.io/fp-lgn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04047 2025-07-31 cs.CV 79%

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Ziyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang, Xiaojian Ma, Yixin Chen, Baoxiong Jia, Wei Liang, Qian Yu, Zhidong Deng, Siyuan Huang, Qing Li

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航) State Key Laboratory of General Artificial Intelligence, BIGAI, China(国家一般人工智能重点实验室, BIGAI, 中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Embodied AI; 3D Vision Language Understanding; ICCV 2025 Highlight; https://mtu3d.github.io; Spatial intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21507 2025-07-30 cs.CV cs.MM 79%

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding

Shibo Gao, Peipei Yang, Yangyang Liu, Yi Chen, Han Zhu, Xuyao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 21 pages, 19 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21450 2025-07-30 cs.CV cs.RO 79%

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

Bolei Chen, Jiaxu Kang, Yifei Wang, Ping Zhong, Qi Wu, Jianxin Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Submitted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21080 2025-07-30 cs.CL cs.AI 79%

Which symbol grounding problem should we try to solve?

Vincent C. Müller

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref (2015) Journal of Experimental and Theoretical Artificial Intelligence, 27 (1, ed. D. Jones & A. Beavers), 73-78

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19847 2025-07-30 cs.CV 79%

Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection

Wenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng, Lei Zhang

机构 * Hong Kong Polytechnic University(香港理工大学) Eastern Institute of Technology(东部技术研究所) Stanford University(斯坦福大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴堡研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20188 2025-07-29 cs.CV 79%

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Mohammed-En-Nadhir Zighem, Abdenour Hadid

机构 * Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,阿布扎比分校,阿联酋)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11261 2025-07-29 cs.CV 79%

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

机构 * South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Singapore Management University(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 79%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01092 2025-07-23 cs.CV cs.RO eess.IV 79%

One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes

Wanjun Jia, Fan Yang, Mengfei Duan, Xianchi Chen, Yinxi Wang, Yiming Jiang, Wenrui Chen, Kailun Yang, Zhiyong Li

机构 * School of Artificial Intelligence and Robotics, Hunan University, China(人工智能与机器人学院,湖南大学) National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(机器人视觉感知与控制技术国家工程研究中心,湖南大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to IROS 2025. Source code and benchmark dataset will be publicly available at https://github.com/Dikay1/OS-AGDO

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14426 2025-07-22 cs.CV 79%

CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding

Zhou Chen, Joe Lin, Sathyanarayanan N. Aakur

机构 * Auburn University(亚伯拉罕大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to NeSy 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01558 2025-07-21 cs.HC cs.AI 79%

Visual Grounding Methods for Efficient Interaction with Desktop Graphical User Interfaces

El Hassane Ettifouri, Jessica López Espejel, Laura Minkova, Tassnim Dardouri, Walid Dahhane

机构 * Research and Innovation Lab, Novelis(创新研究实验室,Novelis)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Preprint submitted to Engineering Applications of Artificial Intelligence journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12232 2025-07-17 cs.CV 79%

MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng

机构 * Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng(作者)

专题命中 视觉定位与Grounding :VLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12123 2025-07-17 cs.CV 79%

Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

Sergey Linok, Gleb Naumov

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 13 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07744 2025-07-11 cs.CV 79%

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

David Pujol-Perich, Sergio Escalera, Albert Clapés

机构 * Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06719 2025-07-10 cs.CV cs.RO 79%

A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding

Zhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao, Longfei Liang, Xiangyang Xue, Yanwei Fu

机构 * Fudan University, Shanghai Innovation Institute(复旦大学) Zhejiang University(浙江大学) Tongji University(同济大学) NeuHelium Co., Ltd(NeuHelium公司) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05424 2025-07-09 cs.CL cs.AI 79%

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models

Yufei Tao, Adam Hiatt, Rahul Seetharaman, Ameeta Agrawal

机构 * Portland State University(波特兰州立大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04741 2025-07-08 cs.CV 79%

Vision-Language Models Can't See the Obvious

Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac, Wamiq Reyaz Para, Ankit Singh, Sanath Narayan

机构 * Technology Innovation Institute(技术创新研究所)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14607 2025-07-01 cs.CV 79%

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang, Wei-Shi Zheng, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学) Southern University of Science and Technology(南方科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Project page: \url{https://isee-laboratory.github.io/ReferDINO}

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22817 2025-07-01 cs.CV 79%

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

Xingyilang Yin, Jiale Wang, Xi Yang, Mutian Xu, Xu Gu, Nannan Wang

机构 * Xidian University(西安电子科技大学) SSE, CUHKSZ(华南理工大学深圳校区)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20384 2025-06-26 cs.AI 79%

Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios

Dror Ivry, Oran Nahum

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18476 2025-06-24 cs.CV 79%

Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding

Yaokun Zhong, Siyu Jiang, Jian Zhu, Jian-Fang Hu

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Guangdong Province Key Laboratory of Information Security Technology, China(广东省信息安全技术重点实验室) Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(教育部机器智能与高级计算重点实验室) Guangdong University of Foreign Studies(广东外语外贸大学) Guangdong University of Technology(广东工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16493 2025-06-23 cs.RO cs.AI 79%

Grounding Language Models with Semantic Digital Twins for Robotic Planning

Mehreen Naeem, Andrew Melnik, Michael Beetz

机构 * Institute for Artificial Intelligence, University of Bremen(人工智能研究所,不莱梅大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14495 2025-06-18 cs.CV 79%

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(都市数据科学部门,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14238 2025-06-18 cs.CV 79%

Unified Representation Space for 3D Visual Grounding

Yinuo Zheng, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(数据科学部,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07896 2025-06-10 cs.AI cs.CL 79%

Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark

Shoko Oka

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 52 pages, Additional resources available on GitHub repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07671 2025-06-10 cs.CL cs.AI 79%

GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation

Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi, Christopher Davis, Adrià de Gispert

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20389 2025-06-10 cs.CV 79%

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang, Ayush Jain, Sriram Yenamandra, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson, Jeong Joon Park, Alexander Sax

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Project page: https://liftgs.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14853 2025-06-10 cs.CV 79%

Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

Peipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng, Zhihua Xia, Chip Hong Chang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏