arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26465 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7494 篇

2502.14638 2025-02-21 cs.CL cs.CV 79%

NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization

Zheyuan Zhang, Runze Li, Tasnim Kabir, Jordan Boyd-Graber

机构 * Tsinghua University(清华大学) Nanjing University(南京大学) University of Maryland(马里兰大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00663 2025-02-20 cs.CV cs.RO 79%

Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment

Kangcheng Liu, Yong-Jin Liu, Baoquan Chen

机构 * California Institute of Technology (Caltech)(加州理工学院) Tsinghua University(清华大学) Peking University(北京大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence, Manuscript Info: 17 Pages, 13 Figures, and 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12917 2025-02-19 cs.CV 79%

Contrast-Unity for Partially-Supervised Temporal Sentence Grounding

Haicheng Wang, Chen Ju, Weixiong Lin, Chaofan Ma, Shuai Xiao, Ya Zhang, Yanfeng Wang

机构 * Taobao & Tmall Group of Alibaba(阿里巴巴淘宝天猫集团) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) CMIC, Shanghai Jiao Tong University(上海交通大学CMIC(媒体与通信工程中心))

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025.The first two authors share the same contribution. arXiv admin note: text overlap with arXiv:2302.09850

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10455 2025-02-18 cs.LG cs.MM 79%

E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection

Junjie Wu, Yumeng Fu, Nan Yu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17207 2025-02-13 cs.CV cs.RO 79%

NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar

Runwei Guan, Jianan Liu, Liye Jia, Haocheng Zhao, Shanliang Yao, Xiaohui Zhu, Ka Lok Man, Eng Gee Lim, Jeremy Smith, Yutao Yue

机构 * University of Liverpool(利物浦大学) Xi’an Jiaotong-Liverpool University(西交利物浦大学) Momoni AI HKUST (GZ)(香港科技大学(广州))

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13382 2025-02-04 cs.CV 79%

VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Yongxin Guo, Jingyu Liu, Mingda Li, Dingxin Cheng, Xiaoying Tang, Dianbo Sui, Qingbin Liu, Xi Chen, Kevin Zhao

机构 * Tencent PCG(腾讯平台与内容事业群)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17654 2025-01-30 cs.CL cs.AI 79%

Exploring Vision Language Models for Multimodal and Multilingual Stance Detection

Jake Vasilakes, Carolina Scarton, Zhixue Zhao

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);分类 cs.AI

Comments Submitted to the International AAAI Conference on Web and Social Media (ICWSM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13925 2025-01-24 cs.CV 79%

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad S. Khan, Salman Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Australian National University(澳大利亚国立大学) The University of Western Australia(西澳大利亚大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09428 2025-01-17 cs.CV 79%

AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring

Xinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo, Xun Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04693 2025-01-16 cs.RO cs.AI 79%

Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

Joshua Jones, Oier Mees, Carmelo Sferrazza, Kyle Stachowicz, Pieter Abbeel, Sergey Levine

机构 * Berkeley AI Research (BAIR)(伯克利人工智能研究院(BAIR)) UC Berkeley(加州大学伯克利分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06746 2025-01-15 cs.CV 79%

Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding

Junlong Ren, Gangjian Zhang, Haifeng Sun, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07396 2025-01-14 cs.CV 79%

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models

Yasiru Ranasinghe, Vibashan VS, James Uplinger, Celso De Melo, Vishal M. Patel

机构 * The Johns Hopkins University(约翰斯·霍普金斯大学) DEVCOM Army Research Laboratory(DEVCOM陆军研究实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05228 2025-01-10 cs.CV 79%

Harnessing Large Language and Vision-Language Models for Robust Out-of-Distribution Detection

Pei-Kang Lee, Jun-Cheng Chen, Ja-Ling Wu

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01416 2025-01-03 cs.CV 79%

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

Yaxian Wang, Henghui Ding, Shuting He, Xudong Jiang, Bifan Wei, Jun Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20059 2024-12-31 cs.CV 79%

AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models

Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Shahid Munir Shah, Mahmoud Aljawarneh, Abdul Akbar Khan, Muhammad Hamzah Siddiqui

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);分类 cs.CV

Comments N-A

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16723 2024-12-24 cs.CV 79%

Divide and Conquer: Grounding a Bleeding Areas in Gastrointestinal Image with Two-Stage Model

Yu-Fan Lin, Bo-Cheng Qiu, Chia-Ming Lee, Chih-Chung Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07870 2024-12-23 cs.CL cs.AI 79%

Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders

Xiaofeng Zhu, Jaya Krishna Mandivarapu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Journal ref EMNLP CustomNLP4U 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22306 2024-12-23 cs.CV 79%

Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial Attention

Haomeng Zhang, Chiao-An Yang, Raymond A. Yeh

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00143 2024-12-20 cs.CV 79%

Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding

Xiaolong Sun, Liushuai Shi, Le Wang, Sanping Zhou, Kun Xia, Yabing Wang, Gang Hua

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by AAAI-25. Code is available at https://github.com/TensorsSun/RGTR

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04925 2024-12-09 cs.CV 79%

$S^3$: Synonymous Semantic Space for Improving Zero-Shot Generalization of Vision-Language Models

Xiaojie Yin, Qilong Wang, Bing Cao, Qinghua Hu

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00890 2024-12-03 cs.CV 79%

Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection

Kun Qian, Tianyu Sun, Wenhong Wang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12066 2024-12-03 cs.CV 79%

Temporally Grounding Instructional Diagrams in Unconstrained Videos

Jiahao Zhang, Frederic Z. Zhang, Cristian Rodriguez, Yizhak Ben-Shabat, Anoop Cherian, Stephen Gould

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to WACV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03077 2024-12-03 cs.CV 79%

MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding

Chun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier Stricker

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19220 2024-12-02 cs.CV cs.MM 79%

Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection

Tsun-Hin Cheung, Ka-Chun Fung, Songjiang Lai, Kwan-Ho Lin, Vincent Ng, Kin-Man Lam

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to APSIPA ASC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01400 2024-12-02 cs.CV 79%

GalLoP: Learning Global and Local Prompts for Vision-Language Models

Marc Lafon, Elias Ramzi, Clément Rambour, Nicolas Audebert, Nicolas Thome

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Journal ref The 18th European Conference on Computer Vision ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17481 2024-11-27 cs.CV 79%

Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding

Mengzhao Wang, Huafeng Li, Yafei Zhang, Jinxing Li, Minghong Xie, Dapeng Tao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This work has been accepted with mandatory minor revisions by TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16932 2024-11-27 cs.CV 79%

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding

Andong Deng, Zhongpai Gao, Anwesa Choudhuri, Benjamin Planche, Meng Zheng, Bin Wang, Terrence Chen, Chen Chen, Ziyan Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08685 2024-11-20 cs.CV 79%

CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding

Linhui Xiao, Xiaoshan Yang, Fang Peng, Ming Yan, Yaowei Wang, Changsheng Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transaction on Multimedia (2023), Paper page: https://ieeexplore.ieee.org/abstract/document/10269126. Code are available at https://github.com/linhuixiao/CLIP-VG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07945 2024-11-13 cs.CV 79%

SimBase: A Simple Baseline for Temporal Video Grounding

Peijun Bao, Alex C. Kot

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02742 2024-11-13 cs.CL cs.LG cs.RO 79%

Grounding Large Language Models In Embodied Environment With Imperfect World Models

Haolan Liu, Jishen Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏