arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2401.02361 2024-01-08 cs.CV 79%

An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Xiangyu Zhao, Yicheng Chen, Shilin Xu, Xiangtai Li, Xinjiang Wang, Yining Li, Haian Huang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01578 2024-01-04 cs.CV 79%

Context-Guided Spatio-Temporal Video Grounding

Xin Gu, Heng Fan, Yan Huang, Tiejian Luo, Libo Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17189 2024-01-01 cs.CV 79%

Exploring Iterative Refinement with Diffusion Models for Video Grounding

Xiao Liang, Tao Shi, Yaoyuan Liang, Te Tao, Shao-Lun Huang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15162 2023-12-27 cs.CV 79%

Cycle-Consistency Learning for Captioning and Grounding

Ning Wang, Jiajun Deng, Mingbo Jia

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments To appear in AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15043 2023-12-27 cs.CV 79%

GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection

Haozhan Shen, Tiancheng Zhao, Mingwei Zhu, Jianwei Yin

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13633 2023-12-22 cs.CV 79%

Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding

Haifeng Huang, Yang Zhao, Zehan Wang, Yan Xia, Zhou Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07207 2023-12-21 cs.CV cs.CL 79%

Beyond Grounding: Extracting Fine-Grained Event Hierarchies Across Modalities

Hammad A. Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh, Long Chen, Yulei Niu, Xudong Lin, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11967 2023-12-20 cs.CV 79%

Context Disentangling and Prototype Inheriting for Robust Visual Grounding

Wei Tang, Liang Li, Xuejing Liu, Lu Jin, Jinhui Tang, Zechao Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09532 2023-12-18 cs.AI 79%

Grounding for Artificial Intelligence

Bing Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08022 2023-12-14 cs.CV 79%

Mono3DVG: 3D Visual Grounding in Monocular Images

Yang Zhan, Yuan Yuan, Zhitong Xiong

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by the Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00541 2023-12-14 cs.CV 79%

Detecting Cloud Presence in Satellite Images Using the RGB-based CLIP Vision-Language Model

Mikolaj Czerkawski, Robert Atkinson, Christos Tachtatzis

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Journal ref IGARSS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04794 2023-12-11 cs.CV 79%

Visual Grounding of Whole Radiology Reports for 3D CT Images

Akimichi Ichinose, Taro Hatsutani, Keigo Nakamura, Yoshiro Kitamura, Satoshi Iizuka, Edgar Simo-Serra, Shoji Kido, Noriyuki Tomiyama

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 14 pages, 7 figures. Accepted at MICCAI 2023

Journal ref Medical Image Computing and Computer Assisted Intervention Lecture Notes in Computer Science 14224 (2023) 611-621

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11887 2023-11-21 cs.CV 79%

A Unified Framework for 3D Point Cloud Visual Grounding

Haojia Lin, Yongdong Luo, Xiawu Zheng, Lijiang Li, Fei Chao, Taisong Jin, Donghao Luo, Yan Wang, Liujuan Cao, Rongrong Ji

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02536 2023-11-07 cs.CV 79%

Augment the Pairs: Semantics-Preserving Image-Caption Pair Augmentation for Grounding-Based Vision and Language Models

Jingru Yi, Burak Uzkent, Oana Ignat, Zili Li, Amanmeet Garg, Xiang Yu, Linda Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to WACV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18773 2023-10-31 cs.CV 79%

CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data

Taiki Miyanishi, Fumiya Kitamori, Shuhei Kurita, Jungdae Lee, Motoaki Kawanabe, Nakamasa Inoue

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments NeurIPS D&B 2023. The first two authors are equally contributed

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18142 2023-10-30 cs.CV 79%

Semi-Supervised Panoptic Narrative Grounding

Danni Yang, Jiayi Ji, Xiaoshuai Sun, Haowei Wang, Yinan Li, Yiwei Ma, Rongrong Ji

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17395 2023-10-27 cs.CV 79%

Learning Temporal Sentence Grounding From Narrated EgoVideos

Kevin Flanagan, Dima Damen, Michael Wray

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted in BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13959 2023-10-27 cs.CV 79%

Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding

Fengyuan Shi, Ruopeng Gao, Weilin Huang, Limin Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in October 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16616 2023-10-26 cs.CV cs.CL 79%

Context Does Matter: End-to-end Panoptic Narrative Grounding with Deformable Attention Refined Matching Network

Yiming Lin, Xiao-Bo Jin, Qiufeng Wang, Kaizhu Huang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICDM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14374 2023-10-24 cs.CV 79%

OV-VG: A Benchmark for Open-Vocabulary Visual Grounding

Chunlei Wang, Wenquan Feng, Xiangtai Li, Guangliang Cheng, Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12147 2023-10-19 cs.RO cs.CV 79%

InViG: Benchmarking Interactive Visual Grounding with 500K Human-Robot Interactions

Hanbo Zhang, Jie Xu, Yuchen Mo, Tao Kong

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 8 pages, 9 figures, 3 tables, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11649 2023-10-19 cs.RO cs.AI cs.CL cs.FL 79%

Grounding Complex Natural Language Commands for Temporal Tasks in Unseen Environments

Jason Xinyu Liu, Ziyi Yang, Ifrah Idrees, Sam Liang, Benjamin Schornstein, Stefanie Tellex, Ankit Shah

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Conference on Robot Learning 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.06272 2023-10-13 cs.CV 79%

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao, Haibing Ren, Hao Shen, Huaxia Xia, Si Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments CVPR 2022, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14941 2023-10-09 cs.CV cs.RO 79%

EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding

Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, Jian Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments CVPR2023, with supplementary material

Journal ref 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 2023, pp. 19231-19242

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15940 2023-09-29 cs.RO cs.CV 79%

Context-Aware Entity Grounding with Open-Vocabulary 3D Scene Graphs

Haonan Chang, Kowndinya Boyalakuntla, Shiyang Lu, Siwei Cai, Eric Jing, Shreesh Keskar, Shijie Geng, Adeeb Abbas, Lifeng Zhou, Kostas Bekris, Abdeslam Boularias

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments The code and dataset used for evaluation can be found at https://github.com/changhaonan/OVSG}{https://github.com/changhaonan/OVSG. This paper has been accepted by CoRL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06135 2023-09-29 cs.RO cs.AI 79%

SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning

Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, Niko Suenderhauf

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted for oral presentation at the Conference on Robot Learning (CoRL), 2023. Project page can be found here: https://sayplan.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14203 2023-09-26 cs.CV 79%

Detecting and Grounding Multi-Modal Media Manipulation and Beyond

Rui Shao, Tianxing Wu, Jianlong Wu, Liqiang Nie, Ziwei Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Extension of our CVPR 2023 paper: arXiv:2304.02556 Code: https://github.com/rshaojimmy/MultiModal-DeepFake

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09667 2023-09-19 cs.CV 79%

Unified Frequency-Assisted Transformer Framework for Detecting and Grounding Multi-Modal Manipulation

Huan Liu, Zichang Tan, Qiang Chen, Yunchao Wei, Yao Zhao, Jingdong Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11319 2023-09-19 cs.RO cs.AI 79%

"Tidy Up the Table": Grounding Common-sense Objective for Tabletop Object Rearrangement

Yiqing Xu, David Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments RSSLRL2023 Workshop, Under review for conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01615 2023-09-19 cs.CV 79%

ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax

Zachary Huemann, Xin Tie, Junjie Hu, Tyler J. Bradshaw

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏