arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2508.05684 2025-08-11 cs.CR cs.LG 79%

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

Junhao He, Tianyu Liu, Jingyuan Zhao, Benjamin Turner

机构 * Huaiyin Institute of Technology(淮阴职业技术学院) Universidad Autónoma de Asunción(阿斯unción自治大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04546 2025-08-07 cs.CV 79%

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

Minghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王宣计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04299 2025-08-07 cs.CV 79%

Length Matters: Length-Aware Transformer for Temporal Sentence Grounding

Yifan Wang, Ziyi Liu, Xiaolong Sun, Jiawei Wang, Hongmin Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03654 2025-08-06 cs.CL cs.CV 79%

Can Large Vision-Language Models Understand Multimodal Sarcasm?

Xinyu Wang, Yue Zhang, Liqiang Jing

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 视觉定位与Grounding :vision-language model(title);visual language model(abstract);分类 cs.CV

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02479 2025-08-05 cs.CV 79%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01699 2025-08-05 cs.CV 79%

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, Song Bai

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01402 2025-08-05 cs.CV 79%

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

Chuangchuang Tan, Jinglu Wang, Xiang Ming, Renshuai Tao, Yunchao Wei, Yao Zhao, Yan Lu

机构 * Beijing Jiaotong University(北京交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04351 2025-08-01 cs.CV 79%

LidaRefer: Context-aware Outdoor 3D Visual Grounding for Autonomous Driving

Yeong-Seung Baek, Heung-Seon Oh

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19947 2025-07-31 cs.RO cs.CL cs.IT cs.LG cs.SY eess.SY math.IT 79%

Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations

Supawich Sitdhipol, Waritwong Sukprasongdee, Ekapol Chuangsuwanich, Rina Tse

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments Accepted to the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC); Supplementary video: https://cu-asl.github.io/fp-lgn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04047 2025-07-31 cs.CV 79%

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Ziyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang, Xiaojian Ma, Yixin Chen, Baoxiong Jia, Wei Liang, Qian Yu, Zhidong Deng, Siyuan Huang, Qing Li

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Beihang University(北航) State Key Laboratory of General Artificial Intelligence, BIGAI, China(国家一般人工智能重点实验室, BIGAI, 中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Embodied AI; 3D Vision Language Understanding; ICCV 2025 Highlight; https://mtu3d.github.io; Spatial intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21507 2025-07-30 cs.CV cs.MM 79%

VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding

Shibo Gao, Peipei Yang, Yangyang Liu, Yi Chen, Han Zhu, Xuyao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 21 pages, 19 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21450 2025-07-30 cs.CV cs.RO 79%

Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

Bolei Chen, Jiaxu Kang, Yifei Wang, Ping Zhong, Qi Wu, Jianxin Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Submitted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21080 2025-07-30 cs.CL cs.AI 79%

Which symbol grounding problem should we try to solve?

Vincent C. Müller

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref (2015) Journal of Experimental and Theoretical Artificial Intelligence, 27 (1, ed. D. Jones & A. Beavers), 73-78

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19847 2025-07-30 cs.CV 79%

Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection

Wenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng, Lei Zhang

机构 * Hong Kong Polytechnic University(香港理工大学) Eastern Institute of Technology(东部技术研究所) Stanford University(斯坦福大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所) Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴堡研究所) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20188 2025-07-29 cs.CV 79%

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Mohammed-En-Nadhir Zighem, Abdenour Hadid

机构 * Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,阿布扎比分校,阿联酋)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11261 2025-07-29 cs.CV 79%

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

机构 * South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室) Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Singapore Management University(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 79%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01092 2025-07-23 cs.CV cs.RO eess.IV 79%

One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes

Wanjun Jia, Fan Yang, Mengfei Duan, Xianchi Chen, Yinxi Wang, Yiming Jiang, Wenrui Chen, Kailun Yang, Zhiyong Li

机构 * School of Artificial Intelligence and Robotics, Hunan University, China(人工智能与机器人学院,湖南大学) National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(机器人视觉感知与控制技术国家工程研究中心,湖南大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to IROS 2025. Source code and benchmark dataset will be publicly available at https://github.com/Dikay1/OS-AGDO

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14426 2025-07-22 cs.CV 79%

CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding

Zhou Chen, Joe Lin, Sathyanarayanan N. Aakur

机构 * Auburn University(亚伯拉罕大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to NeSy 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01558 2025-07-21 cs.HC cs.AI 79%

Visual Grounding Methods for Efficient Interaction with Desktop Graphical User Interfaces

El Hassane Ettifouri, Jessica López Espejel, Laura Minkova, Tassnim Dardouri, Walid Dahhane

机构 * Research and Innovation Lab, Novelis(创新研究实验室,Novelis)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Preprint submitted to Engineering Applications of Artificial Intelligence journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12232 2025-07-17 cs.CV 79%

MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng

机构 * Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng(作者)

专题命中 视觉定位与Grounding :VLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12123 2025-07-17 cs.CV 79%

Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

Sergey Linok, Gleb Naumov

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 13 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07744 2025-07-11 cs.CV 79%

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

David Pujol-Perich, Sergio Escalera, Albert Clapés

机构 * Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06719 2025-07-10 cs.CV cs.RO 79%

A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding

Zhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao, Longfei Liang, Xiangyang Xue, Yanwei Fu

机构 * Fudan University, Shanghai Innovation Institute(复旦大学) Zhejiang University(浙江大学) Tongji University(同济大学) NeuHelium Co., Ltd(NeuHelium公司) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05424 2025-07-09 cs.CL cs.AI 79%

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models

Yufei Tao, Adam Hiatt, Rahul Seetharaman, Ameeta Agrawal

机构 * Portland State University(波特兰州立大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04741 2025-07-08 cs.CV 79%

Vision-Language Models Can't See the Obvious

Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac, Wamiq Reyaz Para, Ankit Singh, Sanath Narayan

机构 * Technology Innovation Institute(技术创新研究所)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14607 2025-07-01 cs.CV 79%

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang, Wei-Shi Zheng, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学) Southern University of Science and Technology(南方科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Project page: \url{https://isee-laboratory.github.io/ReferDINO}

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22817 2025-07-01 cs.CV 79%

Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

Xingyilang Yin, Jiale Wang, Xi Yang, Mutian Xu, Xu Gu, Nannan Wang

机构 * Xidian University(西安电子科技大学) SSE, CUHKSZ(华南理工大学深圳校区)

专题命中 视觉定位与Grounding :VLM(title);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20384 2025-06-26 cs.AI 79%

Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios

Dror Ivry, Oran Nahum

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18476 2025-06-24 cs.CV 79%

Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding

Yaokun Zhong, Siyu Jiang, Jian Zhu, Jian-Fang Hu

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Guangdong Province Key Laboratory of Information Security Technology, China(广东省信息安全技术重点实验室) Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(教育部机器智能与高级计算重点实验室) Guangdong University of Foreign Studies(广东外语外贸大学) Guangdong University of Technology(广东工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏