arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2509.13747 2025-09-18 cs.CV 79%

Improving Generalized Visual Grounding with Instance-aware Joint Learning

Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Lingfeng Yang, Zhenhua Feng, Wankou Yang, Jingdong Wang

机构 * School of Automation, Southeast University(东南大学自动化学院) Baidu Inc.(百度公司) JiangNan University(江南大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in September 2025

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11866 2025-09-16 cs.CV 79%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10278 2025-09-15 cs.CV 79%

Detecting Text Manipulation in Images using Vision Language Models

Vidit Vidit, Pavel Korshunov, Amir Mohammadi, Christophe Ecabert, Ketan Kotwal, Sébastien Marcel

机构 * IDIAP Research Institute(IDIAP研究 institute)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments Accepted in Synthetic Realities and Biometric Security Workshop BMVC-2025. For paper page see https://www.idiap.ch/paper/textvlmdet/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09584 2025-09-12 cs.CV cs.RO 79%

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(新加坡科技研究局) IPAL, CNRS(法国国家科学研究中心IPAL) CerCo, CNRS(法国国家科学研究中心CerCo)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Abstract Paper (Non-Archival) @ ICCV 2025 NeVi Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06291 2025-09-09 cs.CV 79%

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

Jiangnan Xie, Xiaolong Zheng, Liang Zheng

机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06233 2025-09-09 cs.RO cs.CV 79%

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

Tongxuan Tian, Xuhui Kang, Yen-Ling Kuo

机构 * University of Virginia(弗吉尼亚大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Conference on Robot Learning (CoRL) 2025. Project website: https://o3afford.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05154 2025-09-08 eess.IV cs.CV 79%

VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation

Julia Dietlmeier, Oluwabukola Grace Adegboro, Vayangi Ganepola, Claudia Mazo, Noel E. O'Connor

机构 * Insight Research Ireland Centre for Data Analytics, DCU, Dublin, Ireland(爱尔兰洞察研究爱尔兰数据分析中心,都柏林大学,都柏林) Research Ireland Centre for Research Training in Machine Learning, DCU, Dublin, Ireland(爱尔兰研究爱尔兰机器学习研究培训中心,都柏林大学,都柏林) School of Computing, Dublin City University, Dublin, Ireland(计算学院,都柏林城市大学,都柏林)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Medical Imaging with Deep Learning (MIDL 2025) short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04676 2025-09-08 cs.AI cs.HC 79%

An Approach to Grounding AI Model Evaluations in Human-derived Criteria

Sasha Mitts

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 4 figures, 6 pages, presented at CHI 2025 Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02807 2025-09-04 cs.CV 79%

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

Mennatullah Siam

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Work under review in NeurIPS 2025 with the title "Are we using Motion in Referring Segmentation? A Motion-Centric Evaluation"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19493 2025-09-04 cs.CR cs.CV 79%

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, Dongliang Xu

专题命中 视觉定位与Grounding :MLLM(title);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01197 2025-09-04 cs.CV cs.RO 79%

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu

机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) College of Computer Science, Zhejiang University(浙江大学计算机科学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12553 2025-09-04 cs.CV 79%

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality

Yanming Xiu, Tim Scargill, Maria Gorlatova

机构 * Duke University(杜克大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments The paper has been accepted to the 2025 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01401 2025-08-29 cs.CV 79%

Language-to-Space Programming for Training-Free 3D Visual Grounding

Boyu Mi, Hanqing Wang, Tai Wang, Yilun Chen, Jiangmiao Pang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19165 2025-08-27 cs.CV 79%

Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding

Yuzhen Li, Min Liu, Yuan Bian, Xueping Wang, Zhaoyang Li, Gen Li, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(人工智能与机器人学院,湖南大学) College of Information Science and Engineering, Hunan Normal University(信息科学与工程学院,湖南师范大学) School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18372 2025-08-27 cs.CV 79%

OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(戴维森大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM 79%

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11903 2025-08-19 cs.CV 79%

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

Runhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan, Yanjie Dong, Wei Wang, Qi Chen, Xiping Hu

机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) University of Adelaide(阿德莱德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11350 2025-08-18 cs.CV 79%

HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model

Zhenhao Zhang, Hanqing Wang, Xiangyu Zeng, Ziyu Cheng, Jiaxin Liu, Haoyu Yan, Zhirui Liu, Kaiyang Ji, Tianxiang Gui, Ke Hu, Kangyi Chen, Yahao Fan, Mokai Pan

专题命中 视觉定位与Grounding :multimodal large language model(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10976 2025-08-18 cs.AI 79%

Grounding Rule-Based Argumentation Using Datalog

Martin Diller, Sarah Alice Gaggl, Philipp Hanisch, Giuseppina Monterosso, Fritz Rauschenbach

机构 * Logic Programming and Argumentation Group, TU Dresden, Germany(图灵编程与论证组,德累斯顿理工大学,德国) Knowledge-Based Systems Group, TU Dresden, Germany(知识系统组,德累斯顿理工大学,德国) DIMES - University of Calabria, Italy(迪梅斯-卡拉布里亚大学,意大利)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04058 2025-08-18 cs.CV 79%

LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding

Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00372 2025-08-15 cs.CV 79%

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda, Reza Haffari, Peter J. Stuckey, Hamid Rezatofighi

机构 * Monash University(莫纳什大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08633 2025-08-13 cs.AI cs.LO 79%

Diminution: On Reducing the Size of Grounding ASP Programs

HuanYu Yang, Fengming Zhu, YangFan Wu, Jianmin Ji

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14594 2025-08-12 cs.CV 79%

Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems

Qihao Yuan, Kailai Li, Jiaming Zhang

机构 * Bernoulli Institute for Mathematics, Computer Science and Artificial Intelligence, University of Groningen(格罗宁根大学伯努利学院) Computer Vision for Human-Computer Interaction Lab (cv:hci), Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院人机交互计算机视觉实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05684 2025-08-11 cs.CR cs.LG 79%

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

Junhao He, Tianyu Liu, Jingyuan Zhao, Benjamin Turner

机构 * Huaiyin Institute of Technology(淮阴职业技术学院) Universidad Autónoma de Asunción(阿斯unción自治大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04546 2025-08-07 cs.CV 79%

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

Minghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王宣计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04299 2025-08-07 cs.CV 79%

Length Matters: Length-Aware Transformer for Temporal Sentence Grounding

Yifan Wang, Ziyi Liu, Xiaolong Sun, Jiawei Wang, Hongmin Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03654 2025-08-06 cs.CL cs.CV 79%

Can Large Vision-Language Models Understand Multimodal Sarcasm?

Xinyu Wang, Yue Zhang, Liqiang Jing

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 视觉定位与Grounding :vision-language model(title);visual language model(abstract);分类 cs.CV

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02479 2025-08-05 cs.CV 79%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01699 2025-08-05 cs.CV 79%

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, Song Bai

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01402 2025-08-05 cs.CV 79%

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

Chuangchuang Tan, Jinglu Wang, Xiang Ming, Renshuai Tao, Yunchao Wei, Yao Zhao, Yan Lu

机构 * Beijing Jiaotong University(北京交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏