arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2506.16493 2025-06-23 cs.RO cs.AI 79%

Grounding Language Models with Semantic Digital Twins for Robotic Planning

Mehreen Naeem, Andrew Melnik, Michael Beetz

机构 * Institute for Artificial Intelligence, University of Bremen(人工智能研究所,不莱梅大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14495 2025-06-18 cs.CV 79%

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(都市数据科学部门,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14238 2025-06-18 cs.CV 79%

Unified Representation Space for 3D Visual Grounding

Yinuo Zheng, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(数据科学部,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07896 2025-06-10 cs.AI cs.CL 79%

Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark

Shoko Oka

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 52 pages, Additional resources available on GitHub repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07671 2025-06-10 cs.CL cs.AI 79%

GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation

Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi, Christopher Davis, Adrià de Gispert

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20389 2025-06-10 cs.CV 79%

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang, Ayush Jain, Sriram Yenamandra, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson, Jeong Joon Park, Alexander Sax

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Project page: https://liftgs.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14853 2025-06-10 cs.CV 79%

Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

Peipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng, Zhihua Xia, Chip Hong Chang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05890 2025-06-09 cs.CV 79%

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

Yiheng Li, Yang Yang, Zichang Tan, Huan Liu, Weihua Chen, Xu Zhou, Zhen Lei

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算机辅助研究部) School of Computer Science and Engineering, the Faculty of Innovation Engineering, M.U.S.T(慕斯科技大学计算机科学与工程学院) Sangfor Technologies Inc.(Sangfor技术有限公司) Beijing Jiaotong University(北京交通大学) Alibaba Group(阿里巴巴集团)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17433 2025-06-05 cs.AI 79%

MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

Zhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao, Yifan Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, Xian Wu

机构 * The Chinese University of Hong Kong(香港中文大学) University of International Relations(国际关系大学) Jarvis Research Center, Tencent YouTu Lab(腾讯YouTu实验室) Westlake University(西湖大学)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24103 2025-06-02 cs.CV 79%

Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors

Peiran Xu, Yadong Mu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17143 2025-05-28 cs.AI 79%

ASP-based Multi-shot Reasoning via DLV2 with Incremental Grounding

Francesco Calimeri, Giovambattista Ianni, Francesco Pacenza, Simona Perri, Jessica Zangari

机构 * University of Calabria(卡布里亚大学) Gruppo Nazionale Calcolo Scientifico-Istituto Nazionale di Alta Matematica(国家科学计算集团-国家高级数学研究所) DLVSystem Srl(DLVSystem公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Under consideration in Theory and Practice of Logic Programming (TPLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19242 2025-05-27 cs.CV 79%

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

Alaa Dalaq, Muzammil Behzad

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22436 2025-05-27 cs.CV 79%

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

Fuhao Li, Huan Jin, Bin Gao, Liaoyuan Fan, Lihui Jiang, Long Zeng

机构 * Tsinghua University(清华大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12370 2025-05-27 cs.AI 79%

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

Xinbin Yuan, Jian Zhang, Kaixin Li, Zhuoxuan Cai, Lujian Yao, Jie Chen, Enguang Wang, Qibin Hou, Jinwei Chen, Peng-Tao Jiang, Bo Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15332 2025-05-22 cs.CV 79%

Towards Zero-Shot Differential Morphing Attack Detection with Multimodal Large Language Models

Ria Shekhawat, Hailin Li, Raghavendra Ramachandra, Sushma Venkatesh

机构 * Norwegian University of Science and Technology (NTNU)(挪威科学与技术大学) MOBAI AS(MOBAI公司)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted at IEEE International Conference on Automatic Face and Gesture Recognition (FG 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16282 2025-05-21 cs.CV 79%

Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model

Zhaochong An, Guolei Sun, Yun Liu, Runjia Li, Junlin Han, Ender Konukoglu, Serge Belongie

机构 * Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系) Computer Vision Laboratory, ETH Zurich(苏黎世联邦理工学院计算机视觉实验室) College of Computer Science, Nankai University(南开大学计算机学院) Department of Engineering Science, University of Oxford(牛津大学工程科学系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07407 2025-05-21 cs.CV 79%

SemAug: Semantically Meaningful Image Augmentations for Object Detection Through Language Grounding

Morgan Heisler, Amin Banitalebi-Dehkordi, Yong Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09990 2025-05-20 cs.CV 79%

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Long Cheng, Jiafei Duan, Yi Ru Wang, Haoquan Fang, Boyang Li, Yushan Huang, Elvis Wang, Ainaz Eftekhar, Jason Lee, Wentao Yuan, Rose Hendrix, Noah A. Smith, Fei Xia, Dieter Fox, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能 Allen 机构) Anderson Collegiate Vocational Institute(安德森职业学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 Pages, Dataset and code:https://pointarena.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08725 2025-05-14 cs.CV 79%

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou, Dingyuan Zhang, Hongwei Xie, Bing Wang, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车)

专题命中 视觉定位与Grounding :vision-language model(title);grounding(abstract);分类 cs.CV

Comments The dataset and code will be released at https://github.com/zc-zhao/DriveMonkey

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04830 2025-05-14 cs.CL cs.AI 79%

Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents

Jingying Zeng, Hui Liu, Zhenwei Dai, Xianfeng Tang, Chen Luo, Samarth Varshney, Zhen Li, Qi He

机构 * Amazon(亚马逊)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07768 2025-05-13 cs.SE cs.AI cs.CL 79%

Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding

Yifeng Di, Tianyi Zhang

机构 * Purdue University(普渡大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted to ICSE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06557 2025-05-13 cs.CV 79%

Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining

Lu Dong, Haiyu Zhang, Hongjie Zhang, Yifei Huang, Zhen-Hua Ling, Yu Qiao, Limin Wang, Yali Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) State Key Laboratory for Novel Software Technology(国家软件创新科技重点实验室) Nanjing University(南京大学) Shenzhen Key Laboratory of Computer Vision and Pattern Recognition(深圳计算机视觉与模式识别重点实验室) Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments TCSVT 2025, doi at https://ieeexplore.ieee.org/document/10970001

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04965 2025-05-09 cs.CV 79%

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng, Rui Huang, Yepeng Weng, Zhongchao Shi, Gao Huang

机构 * Department of Automation, BNRist, Tsinghua University(自动化系、BNRist、清华大学) AI Lab, Lenovo Research(联想研究院人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01809 2025-05-06 cs.CV 79%

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

Xiaoqi Li, Jiaming Liu, Nuowei Han, Liang Heng, Yandong Guo, Hao Dong, Yang Liu

机构 * School of CS, Peking University(计算机科学系,北京大学) AI2Robotic Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23668 2025-05-02 cs.AI 79%

MolGround: A Benchmark for Molecular Grounding

Jiaxin Wu, Ting Zhang, Rubing Chen, Wengyu Zhang, Chen Jason Zhang, Xiao-Yong Wei, Li Qing

机构 * The Hong Kong Polytechnic University(香港理工大学) Sichuan University(四川大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11564 2025-05-01 cs.RO cs.CL cs.CV cs.HC 79%

Learning 6-DoF Fine-grained Grasp Detection Based on Part Affordance Grounding

Yaoxian Song, Penglei Sun, Piaopiao Jin, Yi Ren, Yu Zheng, Zhixu Li, Xiaowen Chu, Yue Zhang, Tiefeng Li, Jason Gu

机构 * Center for X-Mechanics, School of Aeronautics and Astronautics, Zhejiang University(浙江大学航空航天学院X力学中心) Information Hub, Data Science and Analytics Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心、数据科学与分析方向) Advanced Manufacturing Lab, Huawei Technologies(华为技术先进制造实验室) Research Institute, UBTECH Robotics Inc.(UBTECH机器人研究院) School of Information and School of Smart Governance, Renmin University of China(中国人民大学信息学院和智能治理学院) School of Engineering, Westlake University and the Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖大学工程学院和西湖先进科技研究院) Department of Electrical and Computer Engineering, Dalhousie University(达尔豪斯大学电气与计算机工程系)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 15 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18738 2025-04-29 cs.CV 79%

A Review of 3D Object Detection with Vision-Language Models

Ranjan Sapkota, Konstantinos I Roumeliotis, Rahul Harsha Cheppally, Marco Flores Calero, Manoj Karkee

机构 * Cornell University(康奈尔大学) University of Peloponnese(希腊皮洛斯大学) Kansas State University(堪萨斯州立大学) Universidad de las Fuerzas Armadas(武装力量大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05091 2025-04-28 cs.CV 79%

DCFormer: Efficient 3D Vision-Language Modeling with Decomposed Convolutions

Gorkem Can Ates, Yu Xin, Kuang Gong, Wei Shao

机构 * Department of Medicine, University of Florida(医学系,佛罗里达大学) Department of Electrical and Computer Engineering, University of Florida(电气与计算机工程系,佛罗里达大学) Department of Biomedical Engineering, University of Florida(生物医学工程系,佛罗里达大学) Intelligent Clinical Care Center, University of Florida(智能临床护理中心,佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15545 2025-04-23 eess.IV cs.CV 79%

VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining

Zizhi Chen, Xinyu Zhang, Minghao Han, Yizhou Liu, Ziyun Qian, Weifeng Zhang, Xukun Zhang, Jingwei Wei, Lihua Zhang

机构 * Fudan University(复旦大学) Central South University(中南大学) Harbin Institute of Technology(哈尔滨工业大学) Chinese Academy of Sciences(中国科学院)

专题命中 视觉定位与Grounding :VLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14553 2025-04-22 cs.CV 79%

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection

Weijun Zhuang, Qizhang Li, Xin Li, Ming Liu, Xiaopeng Hong, Feng Gao, Fan Yang, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室) Peking University(北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏