arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2407.01921 2024-07-08 cs.CV 57%

GVDIFF: Grounded Text-to-Video Generation with Diffusion Models

Huanzhang Dou, Ruixiang Li, Wei Su, Xi Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18070 2024-07-02 cs.CV 57%

EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Baoqi Pei, Guo Chen, Jilan Xu, Yuping He, Yicheng Liu, Kanghua Pan, Yifei Huang, Yali Wang, Tong Lu, Limin Wang, Yu Qiao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Champion solutions in the EgoVis CVPR 2024 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19387 2024-07-02 cs.CV 57%

Video Anomaly Detection in 10 Years: A Survey and Outlook

Moshira Abdalla, Sajid Javed, Muaz Al Radi, Anwaar Ulhaq, Naoufel Werghi

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12235 2024-07-02 cs.CV 57%

Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo, Chuchu Han, Xiaonan Huang, Changxin Gao, Yuehuan Wang, Nong Sang

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09069 2024-07-02 cs.CL cs.AI 57%

How Well Do Large Language Models Truly Ground?

Hyunji Lee, Sejune Joo, Chaeeun Kim, Joel Jang, Doyoung Kim, Kyoung-Woon On, Minjoon Seo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments published at NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09858 2024-06-27 cs.CV 57%

Unsupervised Open-Vocabulary Object Localization in Videos

Ke Fan, Zechen Bai, Tianjun Xiao, Dominik Zietlow, Max Horn, Zixu Zhao, Carl-Johann Simon-Gabriel, Mike Zheng Shou, Francesco Locatello, Bernt Schiele, Thomas Brox, Zheng Zhang, Yanwei Fu, Tong He

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICCV 2023; Presented on CVPR 2024 Workshop CORR; Project Page:https://github.com/amazon-science/object-centric-vol

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00690 2024-06-26 cs.CV 57%

Open-vocabulary object 6D pose estimation

Jaime Corsetti, Davide Boscaini, Changjae Oh, Andrea Cavallaro, Fabio Poiesi

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Camera ready version (CVPR 2024, poster highlight). New Oryon version: arXiv:2406.16384

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16143 2024-06-25 cs.CV 57%

Review of Zero-Shot and Few-Shot AI Algorithms in The Medical Domain

Maged Badawi, Mohammedyahia Abushanab, Sheethal Bhat, Andreas Maier

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15764 2024-06-25 cs.CV 57%

TP-DRSeg: Improving Diabetic Retinopathy Lesion Segmentation with Explicit Text-Prompts Assisted SAM

Wenxue Li, Xinyu Xiong, Peng Xia, Lie Ju, Zongyuan Ge

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04655 2024-06-17 cs.LG 57%

Open-Vocabulary Calibration for Fine-tuned CLIP

Shuoyuan Wang, Jindong Wang, Guoqing Wang, Bob Zhang, Kaiyang Zhou, Hongxin Wei

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

Comments Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08463 2024-06-13 cs.CV 57%

Self-supervised Learning of Neural Implicit Feature Fields for Camera Pose Refinement

Maxime Pietrantoni, Gabriela Csurka, Martin Humenberger, Torsten Sattler

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Published in 3DV24 (highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06908 2024-06-12 cs.CV 57%

UVIS: Unsupervised Video Instance Segmentation

Shuaiyi Huang, Saksham Suri, Kamal Gupta, Sai Saketh Rambhatla, Ser-nam Lim, Abhinav Shrivastava

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR2024 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05917 2024-06-11 cs.CV 57%

Point-VOS: Pointing Up Video Object Segmentation

Idil Esen Zulfikar, Sabarinath Mahadevan, Paul Voigtlaender, Bastian Leibe

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to CVPR2024!

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04675 2024-06-10 cs.CV 57%

OVMR: Open-Vocabulary Recognition with Multi-Modal References

Zehong Ma, Shiliang Zhang, Longhui Wei, Qi Tian

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00670 2024-06-07 cs.CV 57%

Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

Yunheng Li, ZhongYu Li, Quansheng Zeng, Qibin Hou, Ming-Ming Cheng

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03428 2024-06-06 cs.LG 57%

HelloFresh: LLM Evaluations on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits

Tim Franzmeyer, Aleksandar Shtedritski, Samuel Albanie, Philip Torr, João F. Henriques, Jakob N. Foerster

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments ACL 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03085 2024-06-06 cs.LG cs.IR 57%

Exploring User Retrieval Integration towards Large Language Models for Cross-Domain Sequential Recommendation

Tingjia Shen, Hao Wang, Jiaqing Zhang, Sirui Zhao, Liangyue Li, Zulong Chen, Defu Lian, Enhong Chen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01170 2024-06-04 cs.CV 57%

Zero-Shot Out-of-Distribution Detection with Outlier Label Exposure

Choubo Ding, Guansong Pang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by IJCNN2024, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00806 2024-06-04 cs.LG 57%

Envisioning Outlier Exposure by Large Language Models for Out-of-Distribution Detection

Chentao Cao, Zhun Zhong, Zhanke Zhou, Yang Liu, Tongliang Liu, Bo Han

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11248 2024-06-04 cs.CV 57%

CoLLaVO: Crayon Large Language and Vision mOdel

Byung-Kwan Lee, Beomchan Park, Chae Won Kim, Yong Man Ro

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments ACL 2024 Findings. Code available: https://github.com/ByungKwanLee/CoLLaVO

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00510 2024-06-04 cs.CV 57%

Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

Jiaming Li, Jiacheng Zhang, Jichang Li, Ge Li, Si Liu, Liang Lin, Guanbin Li

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14455 2024-06-04 cs.CV 57%

TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing

Teng Xu, Jiamin Chen, Peng Chen, Youjia Zhang, Junqing Yu, Wei Yang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20846 2024-06-03 cs.CL cs.AI 57%

Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models

A. Bavaresco, A. Testoni, R. Fernández

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted to the main conference ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20735 2024-06-03 cs.CV 57%

Language Augmentation in CLIP for Improved Anatomy Detection on Multi-modal Medical Images

Mansi Kakkar, Dattesh Shanbhag, Chandan Aladahalli, Gurunath Reddy M

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments $©$ 2024 IEEE. Accepted in 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09989 2024-05-30 cs.CV cs.CL 57%

LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition

Jinyuan Li, Han Li, Di Sun, Jiahao Wang, Wenkun Zhang, Zan Wang, Gang Pan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18304 2024-05-29 cs.CV 57%

Multi-modal Generation via Cross-Modal In-Context Learning

Amandeep Kumar, Muzammal Naseer, Sanath Narayan, Rao Muhammad Anwer, Salman Khan, Hisham Cholakkal

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10123 2024-05-29 cs.CV 57%

AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion

Yitong Jiang, Zhaoyang Zhang, Tianfan Xue, Jinwei Gu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09593 2024-05-27 cs.CV 57%

Renovating Names in Open-Vocabulary Segmentation Benchmarks

Haiwen Huang, Songyou Peng, Dan Zhang, Andreas Geiger

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11265 2024-05-27 cs.CV 57%

Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions

Wenxuan Wang, Yisi Zhang, Xingjian He, Yichen Yan, Zijia Zhao, Xinlong Wang, Jing Liu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments This work has been accepted by ACL 2024 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14966 2024-05-27 cs.AI 57%

Creativity and Markov Decision Processes

Joonas Lahikainen, Nadia M. Ady, Christian Guckelsberger

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 10 pages, full paper at 15th International Conference on Computational Creativity, ICCC'24

详情

展开后加载摘要…

URL PDF HTML 收藏