arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2306.11729 2024-10-16 cs.CV 57%

Dense Video Object Captioning from Disjoint Supervision

Xingyi Zhou, Anurag Arnab, Chen Sun, Cordelia Schmid

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Code is available at https://github.com/google-research/scenic/tree/main/scenic/projects/densevoc

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04428 2024-10-15 stat.ML cs.LG cs.SY eess.SY 57%

Sample-Efficient Linear Representation Learning from Non-IID Non-Isotropic Data

Thomas T. C. K. Zhang, Leonardo F. Toso, James Anderson, Nikolai Matni

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

Comments Appeared at ICLR 2024 (spotlight presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01429 2024-10-14 cs.CV 57%

EAGLE: Efficient Adaptive Geometry-based Learning in Cross-view Understanding

Thanh-Dat Truong, Utsav Prabhu, Dongyi Wang, Bhiksha Raj, Susan Gauch, Jeyamkondan Subbiah, Khoa Luu

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07331 2024-10-14 cs.CL cs.AI 57%

DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models

Yiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang, Fangyu Lei, Yifan Wei, Shizhu He, Lifu Huang, Xiao Liu, Jun Zhao, Kang Liu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01460 2024-10-11 cs.LG 57%

ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive Advantages

Andrew Jesson, Chris Lu, Gunshi Gupta, Nicolas Beltran-Velez, Angelos Filos, Jakob Nicolaus Foerster, Yarin Gal

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06626 2024-10-10 cs.CV 57%

Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments

Meng Yu, Luojie Yang, Xunjie He, Yi Yang, Yufeng Yue

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05980 2024-10-10 cs.LG 57%

Generalizing to any diverse distribution: uniformity, gentle finetuning and rebalancing

Andreas Loukas, Karolis Martinkus, Ed Wagstaff, Kyunghyun Cho

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05650 2024-10-10 cs.CV cs.MM 57%

SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection

Zishuo Wang, Wenhao Zhou, Jinglin Xu, Yuxin Peng

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05239 2024-10-10 cs.CV cs.CL 57%

TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models

Rabin Adhikari, Safal Thapaliya, Manish Dhakal, Bishesh Khanal

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted at ACCV 2024 (oral presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05008 2024-10-10 cs.CV 57%

FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow

Hangyu Li, Xiangxiang Chu, Dingyuan Shi, Wang Lin

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16032 2024-10-10 cs.LG 57%

Studying Large Language Model Behaviors Under Context-Memory Conflicts With Real Documents

Evgenii Kortukov, Alexander Rubinstein, Elisa Nguyen, Seong Joon Oh

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12455 2024-10-10 cs.CV 57%

CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation

Wenqi Zhu, Jiale Cao, Jin Xie, Shuangming Yang, Yanwei Pang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03900 2024-10-08 cs.CV 57%

The Wallpaper is Ugly: Indoor Localization using Vision and Language

Seth Pate, Lawson L. S. Wong

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments RO-MAN 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02365 2024-10-04 cs.CL cs.AI 57%

From Concrete to Abstract: A Multimodal Generative Approach to Abstract Concept Learning

Haodong Xie, Rahul Singh Maharjan, Federico Tavella, Angelo Cangelosi

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04109 2024-10-02 cs.CV 57%

From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models

Changming Xiao, Qi Yang, Feng Zhou, Changshui Zhang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments A revised version of this paper will be published in Neurocomputing, see https://doi.org/10.1016/j.neucom.2024.128437

Journal ref Neurocomputing, Volume 610, 2024, 128437

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20259 2024-10-01 cs.AI 57%

Learning to Ground Existentially Quantified Goals

Martin Funkquist, Simon Ståhlberg, Hector Geffner

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 11 pages, Accepted at the 21st International Conference on Principles of Knowledge Representation and Reasoning (KR2024) in the Reasoning, Learning, and Decision Making track

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19846 2024-10-01 cs.CV 57%

Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels

Heeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho, Anurag Arnab, Paul Hongsuck Seo, Seungryong Kim

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments To appear at NeurIPS 2024. Project page is available at https://cvlab-kaist.github.io/PixelCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19816 2024-10-01 cs.RO cs.AI 57%

Grounded Curriculum Learning

Linji Wang, Zifan Xu, Peter Stone, Xuesu Xiao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18372 2024-10-01 cs.CV 57%

You Only Speak Once to See

Wenhao Yang, Jianguo Wei, Wenhuan Lu, Lei Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09316 2024-10-01 cs.CV 57%

Diffusion Models for Open-Vocabulary Segmentation

Laurynas Karazija, Iro Laina, Andrea Vedaldi, Christian Rupprecht

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18111 2024-09-27 cs.CV 57%

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Ye Liu, Zongyang Ma, Zhongang Qi, Yang Wu, Ying Shan, Chang Wen Chen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2024 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17580 2024-09-27 cs.IR cs.AI cs.DB 57%

Enhancing Structured-Data Retrieval with GraphRAG: Soccer Data Case Study

Zahra Sepasdar, Sushant Gautam, Cise Midoglu, Michael A. Riegler, Pål Halvorsen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16159 2024-09-25 cs.CV 57%

ComiCap: A VLMs pipeline for dense captioning of Comic Panels

Emanuele Vivoli, Niccolò Biondi, Marco Bertini, Dimosthenis Karatzas

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop (AI for Visual Art), repo: https://github.com/emanuelevivoli/ComiCap

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16036 2024-09-25 q-bio.NC cs.AI 57%

Grounded Computation & Consciousness: A Framework for Exploring Consciousness in Machines & Other Organisms

Ryan Williams

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04449 2024-09-23 cs.CL cs.CV 57%

MAIRA-2: Grounded Radiology Report Generation

Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Anton Schwaighofer, Anja Thieme, Sam Bond-Taylor, Maximilian Ilse, Fernando Pérez-García, Valentina Salvatelli, Harshita Sharma, Felix Meissen, Mercy Ranjit, Shaury Srivastav, Julia Gong, Noel C. F. Codella, Fabian Falck, Ozan Oktay, Matthew P. Lungren, Maria Teodora Wetscherek, Javier Alvarez-Valle, Stephanie L. Hyland

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 72 pages, 21 figures. v2 updates the model and adds results on the PadChest-GR dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13162 2024-09-23 cs.CV 57%

Towards Zero-shot Point Cloud Anomaly Detection: A Multi-View Projection Framework

Yuqi Cheng, Yunkang Cao, Guoyang Xie, Zhichao Lu, Weiming Shen

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12580 2024-09-20 cs.CV 57%

LLMs Can Check Their Own Results to Mitigate Hallucinations in Traffic Understanding Tasks

Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

Comments ICTSS 2024, 36th International Conference on Testing Software and Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10917 2024-09-18 cs.CV 57%

AMEGO: Active Memory from long EGOcentric videos

Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, Dima Damen

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted to ECCV 2024. Project webpage: https://gabrielegoletto.github.io/AMEGO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07448 2024-09-18 cs.CV cs.CL eess.IV 57%

Transferable and Principled Efficiency for Open-Vocabulary Segmentation

Jingxuan Xu, Wuyang Chen, Yao Zhao, Yunchao Wei

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11322 2024-09-17 cs.CL cs.AI 57%

StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows

Yiran Wu, Tianwei Yue, Shaokun Zhang, Chi Wang, Qingyun Wu

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏