arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7473 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2411.19220 2024-12-02 cs.CV cs.MM 79%

Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection

Tsun-Hin Cheung, Ka-Chun Fung, Songjiang Lai, Kwan-Ho Lin, Vincent Ng, Kin-Man Lam

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to APSIPA ASC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01400 2024-12-02 cs.CV 79%

GalLoP: Learning Global and Local Prompts for Vision-Language Models

Marc Lafon, Elias Ramzi, Clément Rambour, Nicolas Audebert, Nicolas Thome

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Journal ref The 18th European Conference on Computer Vision ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17481 2024-11-27 cs.CV 79%

Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding

Mengzhao Wang, Huafeng Li, Yafei Zhang, Jinxing Li, Minghong Xie, Dapeng Tao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This work has been accepted with mandatory minor revisions by TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16932 2024-11-27 cs.CV 79%

Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding

Andong Deng, Zhongpai Gao, Anwesa Choudhuri, Benjamin Planche, Meng Zheng, Bin Wang, Terrence Chen, Chen Chen, Ziyan Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08685 2024-11-20 cs.CV 79%

CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding

Linhui Xiao, Xiaoshan Yang, Fang Peng, Ming Yan, Yaowei Wang, Changsheng Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transaction on Multimedia (2023), Paper page: https://ieeexplore.ieee.org/abstract/document/10269126. Code are available at https://github.com/linhuixiao/CLIP-VG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07945 2024-11-13 cs.CV 79%

SimBase: A Simple Baseline for Temporal Video Grounding

Peijun Bao, Alex C. Kot

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02742 2024-11-13 cs.CL cs.LG cs.RO 79%

Grounding Large Language Models In Embodied Environment With Imperfect World Models

Haolan Liu, Jishen Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03409 2024-11-07 cs.RO cs.AI 79%

STEER: Flexible Robotic Manipulation via Dense Language Grounding

Laura Smith, Alex Irpan, Montserrat Gonzalez Arenas, Sean Kirmani, Dmitry Kalashnikov, Dhruv Shah, Ted Xiao

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);分类 cs.AI

Comments Project website: https://lauramsmith.github.io/steer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03405 2024-11-07 cs.CV 79%

Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding

Sombit Dey, Ozan Unal, Christos Sakaridis, Luc Van Gool

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00881 2024-11-05 cs.CV 79%

Technical Report for SoccerNet Challenge 2022 -- Replay Grounding Task

Shimin Chen, Wei Li, Jiaming Chu, Chen Chen, Chen Zhang, Yandong Guo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20474 2024-11-04 cs.CV 79%

GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

Phillip Y. Lee, Taehoon Yoon, Minhyuk Sung

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2024. Project Page: https://groundit-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23570 2024-11-01 cs.CV 79%

Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding

Minghong Xie, Mengzhao Wang, Huafeng Li, Yafei Zhang, Dapeng Tao, Zhengtao Yu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This work has been accepted by TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19137 2024-11-01 cs.CV 79%

CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models

Saurav Jha, Dong Gong, Lina Yao

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted as a poster at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12657 2024-10-28 cs.CV 79%

Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding

Jiazhen Wang, Bin Liu, Changtao Miao, Zhiwei Zhao, Wanyi Zhuang, Qi Chu, Nenghai Yu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication. Camera-ready version and supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17379 2024-10-23 cs.RO cs.AI 79%

EMPOWER: Embodied Multi-role Open-vocabulary Planning with Online Grounding and Execution

Francesco Argenziano, Michele Brienza, Vincenzo Suriani, Daniele Nardi, Domenico D. Bloisi

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted at IROS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15615 2024-10-22 cs.CV 79%

Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding

Yang Liu, Daizong Liu, Wei Hu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14544 2024-10-21 cs.AI 79%

Computational Grounding of Responsibility Attribution and Anticipation in LTLf

Giuseppe De Giacomo, Emiliano Lorini, Timothy Parker, Gianmarco Parretti

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13598 2024-10-18 cs.CV 79%

Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding

Jongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae Won Cho, Joon Son Chung

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACMMM 24

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12813 2024-10-18 cs.MM cs.CV 79%

ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models

Mengxue Qu, Xiaodong Chen, Wu Liu, Alicia Li, Yao Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12369 2024-10-17 cs.CV 79%

Context-Infused Visual Grounding for Art

Selina Khan, Nanne van Noord

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12225 2024-10-17 cs.CV 79%

Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety

Lucas Choi, Ross Greer

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14494 2024-10-15 cs.CV 79%

Revisiting Few-Shot Object Detection with Vision-Language Models

Anish Madan, Neehar Peri, Shu Kong, Deva Ramanan

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments The first two authors contributed equally. This work has been accepted to the Neural Information Processing Systems (NeurIPS) 2024 Datasets & Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16475 2024-10-11 cs.CV 79%

Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models

Yu-Wei Zhan, Fan Liu, Xin Luo, Xin-Shun Xu, Liqiang Nie, Mohan Kankanhalli

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06108 2024-10-10 cs.AI 79%

ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution

Corban Rivera, Grayson Byrd, William Paul, Tyler Feldman, Meghan Booker, Emma Holmes, David Handelman, Bethany Kemp, Andrew Badger, Aurora Schmidt, Krishna Murthy Jatavallabhula, Celso M de Melo, Lalithkumar Seenivasan, Mathias Unberath, Rama Chellappa

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04426 2024-10-08 cs.CV 79%

CoVLM: Leveraging Consensus from Vision-Language Models for Semi-supervised Multi-modal Fake News Detection

Devank, Jayateja Kalla, Soma Biswas

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted in ACCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06214 2024-10-08 cs.CV 79%

CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding

Eslam Abdelrahman, Mohamed Ayman, Mahmoud Ahmed, Habib Slim, Mohamed Elhoseiny

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03161 2024-10-07 cs.AI 79%

Adaptive Masking Enhances Visual Grounding

Sen Jia, Lei Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Code will be available at https://github.com/git-lenny/IMAGE

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10774 2024-10-02 cs.CL cs.AI 79%

MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents

Liyan Tang, Philippe Laban, Greg Durrett

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17330 2024-09-27 cs.CV 79%

VL4AD: Vision-Language Models Improve Pixel-wise Anomaly Detection

Liangyu Zhong, Joachim Sicking, Fabian Hüger, Hanno Gottschalk

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 27 pages, 9 figures, to be published in ECCV 2024 2nd Workshop on Vision-Centric Autonomous Driving (VCAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19071 2024-09-18 cs.CL cs.AI 79%

EmPO: Emotion Grounding for Empathetic Response Generation through Preference Optimization

Ondrej Sotolar, Vojtech Formanek, Alok Debnath, Allison Lahnala, Charles Welch, Lucie FLek

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments v02, 8 pages long paper, EMNLP ACL style

详情

展开后加载摘要…

URL PDF HTML 收藏