arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2509.19096 2025-09-26 cs.CV cs.SE 79%

Investigating Traffic Accident Detection Using Multimodal Large Language Models

Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig

机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) Virtual Vehicle Research GmbH(虚拟车辆研究公司) Control Systems Group (Dept.-E)(控制系统组) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨技术大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24372 2025-09-26 cs.CV 79%

Beyond Quantity: Distribution-Aware Labeling for Visual Grounding

Yichi Zhang, Gongwei Chen, Jun Zhu, Jia Wan, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 18pages, 8figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21188 2025-09-23 cs.CV 79%

GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding

Zijun Lin, Shuting He, Cheston Tan, Bihan Wen

机构 * Nanyang Technological University(南洋理工大学) Centre for Frontier AI Research, A*STAR(前沿人工智能研究中心,A*STAR) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(计算与经济学交叉研究关键实验室,上海财经大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16806 2025-09-23 cs.CV 79%

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

Fan Yang, Yousong Zhu, Xin Li, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Science(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16534 2025-09-23 cs.CL cs.AI 79%

InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding

Cheng Jiayang, Qianqian Zhuang, Haoran Li, Chunkit Chan, Xin Liu, Lin Qiu, Yangqiu Song

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai Jiaotong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 79%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13836 2025-09-18 cs.CV cs.CL 79%

Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models

Weihang Wang, Xinhao Li, Ziyue Wang, Yan Pang, Jielei Zhang, Peiyi Li, Qiang Zhang, Longwen Gao

机构 * Bilibili(哔哩哔哩) UESTC University of Virginia(弗吉尼亚大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by EMNLP2025 Finding

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13747 2025-09-18 cs.CV 79%

Improving Generalized Visual Grounding with Instance-aware Joint Learning

Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Lingfeng Yang, Zhenhua Feng, Wankou Yang, Jingdong Wang

机构 * School of Automation, Southeast University(东南大学自动化学院) Baidu Inc.(百度公司) JiangNan University(江南大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in September 2025

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11866 2025-09-16 cs.CV 79%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10278 2025-09-15 cs.CV 79%

Detecting Text Manipulation in Images using Vision Language Models

Vidit Vidit, Pavel Korshunov, Amir Mohammadi, Christophe Ecabert, Ketan Kotwal, Sébastien Marcel

机构 * IDIAP Research Institute(IDIAP研究 institute)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments Accepted in Synthetic Realities and Biometric Security Workshop BMVC-2025. For paper page see https://www.idiap.ch/paper/textvlmdet/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09584 2025-09-12 cs.CV cs.RO 79%

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(新加坡科技研究局) IPAL, CNRS(法国国家科学研究中心IPAL) CerCo, CNRS(法国国家科学研究中心CerCo)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Abstract Paper (Non-Archival) @ ICCV 2025 NeVi Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06291 2025-09-09 cs.CV 79%

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

Jiangnan Xie, Xiaolong Zheng, Liang Zheng

机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06233 2025-09-09 cs.RO cs.CV 79%

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

Tongxuan Tian, Xuhui Kang, Yen-Ling Kuo

机构 * University of Virginia(弗吉尼亚大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Conference on Robot Learning (CoRL) 2025. Project website: https://o3afford.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05154 2025-09-08 eess.IV cs.CV 79%

VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation

Julia Dietlmeier, Oluwabukola Grace Adegboro, Vayangi Ganepola, Claudia Mazo, Noel E. O'Connor

机构 * Insight Research Ireland Centre for Data Analytics, DCU, Dublin, Ireland(爱尔兰洞察研究爱尔兰数据分析中心,都柏林大学,都柏林) Research Ireland Centre for Research Training in Machine Learning, DCU, Dublin, Ireland(爱尔兰研究爱尔兰机器学习研究培训中心,都柏林大学,都柏林) School of Computing, Dublin City University, Dublin, Ireland(计算学院,都柏林城市大学,都柏林)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Medical Imaging with Deep Learning (MIDL 2025) short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04676 2025-09-08 cs.AI cs.HC 79%

An Approach to Grounding AI Model Evaluations in Human-derived Criteria

Sasha Mitts

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 4 figures, 6 pages, presented at CHI 2025 Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02807 2025-09-04 cs.CV 79%

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

Mennatullah Siam

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Work under review in NeurIPS 2025 with the title "Are we using Motion in Referring Segmentation? A Motion-Centric Evaluation"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19493 2025-09-04 cs.CR cs.CV 79%

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu

专题命中 视觉定位与Grounding :MLLM(title);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01197 2025-09-04 cs.CV cs.RO 79%

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu

机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) College of Computer Science, Zhejiang University(浙江大学计算机科学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12553 2025-09-04 cs.CV 79%

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality

Yanming Xiu, Tim Scargill, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01401 2025-08-29 cs.CV 79%

Language-to-Space Programming for Training-Free 3D Visual Grounding

Boyu Mi, Hanqing Wang, Tai Wang, Yilun Chen, Jiangmiao Pang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19165 2025-08-27 cs.CV 79%

Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding

Yuzhen Li, Min Liu, Yuan Bian, Xueping Wang, Zhaoyang Li, Gen Li, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(人工智能与机器人学院,湖南大学) College of Information Science and Engineering, Hunan Normal University(信息科学与工程学院,湖南师范大学) School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18372 2025-08-27 cs.CV 79%

OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(戴维森大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 79%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11903 2025-08-19 cs.CV 79%

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

Runhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan, Yanjie Dong, Wei Wang, Qi Chen, Xiping Hu

机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) University of Adelaide(阿德莱德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11350 2025-08-18 cs.CV 79%

HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model

Zhenhao Zhang, Hanqing Wang, Xiangyu Zeng, Ziyu Cheng, Jiaxin Liu, Haoyu Yan, Zhirui Liu, Kaiyang Ji, Tianxiang Gui, Ke Hu, Kangyi Chen, Yahao Fan, Mokai Pan

专题命中 视觉定位与Grounding :multimodal large language model(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10976 2025-08-18 cs.AI 79%

Grounding Rule-Based Argumentation Using Datalog

Martin Diller, Sarah Alice Gaggl, Philipp Hanisch, Giuseppina Monterosso, Fritz Rauschenbach

机构 * Logic Programming and Argumentation Group, TU Dresden, Germany(图灵编程与论证组,德累斯顿理工大学,德国) Knowledge-Based Systems Group, TU Dresden, Germany(知识系统组,德累斯顿理工大学,德国) DIMES - University of Calabria, Italy(迪梅斯-卡拉布里亚大学,意大利)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04058 2025-08-18 cs.CV 79%

LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding

Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00372 2025-08-15 cs.CV 79%

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda, Reza Haffari, Peter J. Stuckey, Hamid Rezatofighi

机构 * Monash University(莫纳什大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08633 2025-08-13 cs.AI cs.LO 79%

Diminution: On Reducing the Size of Grounding ASP Programs

HuanYu Yang, Fengming Zhu, YangFan Wu, Jianmin Ji

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14594 2025-08-12 cs.CV 79%

Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems

Qihao Yuan, Kailai Li, Jiaming Zhang

机构 * Bernoulli Institute for Mathematics, Computer Science and Artificial Intelligence, University of Groningen(格罗宁根大学伯努利学院) Computer Vision for Human-Computer Interaction Lab (cv:hci), Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院人机交互计算机视觉实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏