arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2411.09273 2024-11-15 cs.CL cs.AI 74%

Cross-Modal Consistency in Multimodal Large Language Models

Xiang Zhang, Senyu Li, Ning Shi, Bradley Hauer, Zijun Wu, Grzegorz Kondrak, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04892 2024-11-08 cs.CV 74%

In the Era of Prompt Learning with Vision-Language Models

Ankit Jha

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments ICVGIP 2024, Young Faculty Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03491 2024-11-07 cs.CV 74%

An Application-Agnostic Automatic Target Recognition System Using Vision Language Models

Anthony Palladino, Dana Gajewski, Abigail Aronica, Patryk Deptula, Alexander Hamme, Seiyoung C. Lee, Jeff Muri, Todd Nelling, Michael A. Riley, Brian Wong, Margaret Duff

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments Accepted to the Thirty-Seventh Annual Conference on Innovative Applications of Artificial Intelligence (IAAI-25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14509 2024-10-21 cs.CV 74%

CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection

Andrea Appiani, Cigdem Beyan

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07540 2024-10-11 cs.CV 74%

CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection

Guankun Wang, Han Xiao, Huxin Gao, Renrui Zhang, Long Bai, Xiaoxiao Yang, Zhen Li, Hongsheng Li, Hongliang Ren

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05601 2024-10-10 cs.CV 74%

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

Hang Guo, Tao Dai, Zhihao Ouyang, Taolin Zhang, Yaohua Zha, Bin Chen, Shu-tao Xia

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05267 2024-10-08 cs.CL cs.CV 74%

Grounding Partially-Defined Events in Multimodal Data

Kate Sanders, Reno Kriz, David Etter, Hannah Recknor, Alexander Martin, Cameron Carpenter, Jingyang Lin, Benjamin Van Durme

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Preprint; 9 pages; 2024 EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12821 2024-10-07 cs.CL cs.LG 74%

Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy

Liyan Xu, Zhenlin Su, Mo Yu, Jin Xu, Jinho D. Choi, Jie Zhou, Fei Liu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

Comments Accepted to EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14109 2024-09-27 cs.CV 74%

Vision-Language Models Assisted Unsupervised Video Anomaly Detection

Yalong Jiang, Liquan Mao

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07048 2024-09-12 cs.CV 74%

Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations

Keumgang Cha, Donggeun Yu, Junghoon Seo

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments This study was primarily conducted during the latter half of 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11286 2024-08-23 cs.CV 74%

Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model

Mengying Ge, Dongkai Tang, Mingyang Li

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05941 2024-08-13 cs.CR cs.AI 74%

Multimodal Large Language Models for Phishing Webpage Detection and Identification

Jehyun Lee, Peiyuan Lim, Bryan Hooi, Dinil Mon Divakaran

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

Comments To appear in eCrime 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18695 2024-07-29 cs.CV cs.CL 74%

Grounding Language Models for Visual Entity Recognition

Zilin Xiao, Ming Gong, Paola Cascante-Bonilla, Xingyao Zhang, Jie Wu, Vicente Ordonez

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15408 2024-07-23 cs.CV 74%

Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models

Kent Fujiwara, Mikihiro Tanaka, Qing Yu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments To appear at ECCV 2024. Project page: https://kfworks.com/CAR-WP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05363 2024-07-11 cs.CV 74%

Multi-branch Collaborative Learning Network for 3D Visual Grounding

Zhipeng Qian, Yiwei Ma, Zhekai Lin, Jiayi Ji, Xiawu Zheng, Xiaoshuai Sun, Rongrong Ji

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13245 2024-06-25 cs.RO cs.AI cs.CL 74%

A Survey of Robotic Language Grounding: Tradeoffs between Symbols and Embeddings

Vanya Cohen, Jason Xinyu Liu, Raymond Mooney, Stefanie Tellex, David Watkins

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments IJCAI 2024 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05295 2024-06-19 cs.AI cs.FL 74%

Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception

Yunhao Yang, Cyrus Neary, Ufuk Topcu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Accepted as full paper in AAMAS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09756 2024-06-17 cs.CV 74%

Grounding Image Matching in 3D with MASt3R

Vincent Leroy, Yohann Cabon, Jérôme Revaud

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20002 2024-06-10 cs.CV 74%

Grounding and Enhancing Grid-based Models for Neural Fields

Zelin Zhao, Fenglei Fan, Wenlong Liao, Junchi Yan

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments CVPR24 Oral & Best Paper Award Candidate. Pre-rebuttal scores: 555. Post-rebuttal scores: 555

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15961 2024-05-28 cs.CV 74%

Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images

Yiran Luo, Joshua Feinglass, Tejas Gokhale, Kuan-Cheng Lee, Chitta Baral, Yezhou Yang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted at the 3rd CVPR Workshop on Vision Datasets Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13257 2024-03-27 cs.CL cs.AI 74%

Visual Grounding Helps Learn Word Meanings in Low-Data Regimes

Chengxu Zhuang, Evelina Fedorenko, Jacob Andreas

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Accepted by NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08073 2024-03-18 cs.LG cs.PL cs.SE 74%

Grounding Data Science Code Generation with Input-Output Specifications

Yeming Wen, Pengcheng Yin, Kensen Shi, Henryk Michalewski, Swarat Chaudhuri, Alex Polozov

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09493 2024-03-15 cs.CV 74%

Anomaly Detection by Adapting a pre-trained Vision Language Model

Yuxuan Cai, Xinwei He, Dingkang Liang, Ao Tong, Xiang Bai

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08345 2024-02-12 cs.CL cs.AI 74%

Data Distribution Bottlenecks in Grounding Language Models to Knowledge Bases

Yiheng Shu, Zhiwei Yu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11681 2023-12-18 cs.CV cs.MM 74%

VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly Detection

Peng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang, Yanning Zhang

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments Accept to AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03726 2023-12-08 cs.CL cs.AI 74%

Interpretation modeling: Social grounding of sentences by reasoning over their implicit moral judgments

Liesbeth Allein, Maria Mihaela Truşcǎ, Marie-Francine Moens

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11252 2023-09-21 cs.CL cs.CV 74%

The Scenario Refiner: Grounding subjects in images at the morphological level

Claudia Tagliaferri, Sofia Axioti, Albert Gatt, Denis Paperno

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments presented at the LIMO workshop (Linguistic Insights from and for Multimodal Language Processing @KONVENS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15786 2023-07-27 cs.CV 74%

HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language Models

Shan Ning, Longtian Qiu, Yongfei Liu, Xuming He

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments CVPR 2023.Open sourced, Code and Model Available

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12513 2023-06-12 cs.CV 74%

Learning Point-Language Hierarchical Alignment for 3D Visual Grounding

Jiaming Chen, Weixin Luo, Ran Song, Xiaolin Wei, Lin Ma, Wei Zhang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Champion on ECCV 2022 ScanRefer Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04358 2023-04-11 cs.CL cs.AI 74%

WebBrain: Learning to Generate Factually Correct Articles for Queries by Grounding on Large Web Corpus

Hongjing Qian, Yutao Zhu, Zhicheng Dou, Haoqi Gu, Xinyu Zhang, Zheng Liu, Ruofei Lai, Zhao Cao, Jian-Yun Nie, Ji-Rong Wen

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Codes in https://github.com/qhjqhj00/WebBrain

详情

展开后加载摘要…

URL PDF HTML 收藏