arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2210.04150 2023-04-04 cs.CV cs.LG 62%

Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP

Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, Diana Marculescu

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments CVPR 2023. Project page: https://jeff-liangf.github.io/projects/ovseg

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07632 2023-03-29 cs.CV cs.LG 62%

CoReS: Compatible Representations via Stationarity

Niccolo Biondi, Federico Pernici, Matteo Bruni, Alberto Del Bimbo

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. Code: https://github.com/NiccoBiondi/cores-compatibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14814 2023-03-28 cs.CV cs.AI cs.CL 62%

WinCLIP: Zero-/Few-Shot Anomaly Classification and Segmentation

Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, Onkar Dabeer

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to Conference on Computer Vision and Pattern Recognition (CVPR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13040 2023-03-24 cs.CV cs.AI 62%

Open-Vocabulary Object Detection using Pseudo Caption Labels

Han-Cheol Cho, Won Young Jhoo, Wooyoung Kang, Byungseok Roh

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12242 2023-03-23 cs.CV cs.AI 62%

Side Adapter Network for Open-Vocabulary Semantic Segmentation

Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, Xiang Bai

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments CVPR2023 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12914 2023-03-10 cs.CV cs.LG 62%

Open-vocabulary Attribute Detection

María A. Bravo, Sudhanshu Mittal, Simon Ging, Thomas Brox

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

Comments Accepted at CVPR 2023. https://ovad-benchmark.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10687 2023-03-10 cs.LG cs.AI 62%

The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types

Gaurav R. Ghosal, Matthew Zurek, Daniel S. Brown, Anca D. Dragan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Published at AAAI 2023; 10 pages, 5 figures plus appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10978 2023-02-23 cs.CL cs.AI cs.IR cs.LG 62%

Learning to Retrieve Engaging Follow-Up Queries

Christopher Richardson, Sudipta Kar, Anjishnu Kumar, Anand Ramachandran, Omar Zia Khan, Zeynab Raeesy, Abhinav Sethy

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10281 2023-02-22 cs.CV cs.AI cs.CL 62%

LiT Tuned Models for Efficient Species Detection

Andre Nakkab, Benjamin Feuer, Chinmay Hegde

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 5 figures, 1 table, presented at AAAI 2023 conference for the AIAFS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01206 2023-02-09 cs.CL cs.AI cs.LG 62%

WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Shunyu Yao, Howard Chen, John Yang, Karthik Narasimhan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Project page with code, data, demos: https://webshop-pnlp.github.io. v3 is NeurIPS camera ready version. v4 fixes the choice oracle result as per https://github.com/princeton-nlp/WebShop/issues/15

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00275 2023-02-02 cs.CV cs.LG 62%

Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization

Lukas Haas, Silas Alberti, Michal Skreta

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03344 2023-01-10 cs.CL cs.AI cs.CV 62%

Universal Multimodal Representation for Language Understanding

Zhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11876 2022-12-01 cs.CV cs.AI 62%

Open-Vocabulary DETR with Conditional Matching

Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, Chen Change Loy

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

Comments ECCV 2022 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10902 2022-11-24 cs.LG cs.AI cs.FL 62%

Noisy Symbolic Abstractions for Deep RL: A case study with Reward Machines

Andrew C. Li, Zizhao Chen, Pashootan Vaezipoor, Toryn Q. Klassen, Rodrigo Toro Icarte, Sheila A. McIlraith

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments NeurIPS Deep Reinforcement Learning Workshop 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11572 2022-11-22 cs.CV cs.CL cs.LG 62%

Detect Only What You Specify : Object Detection with Linguistic Target

Moyuru Yamada

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13611 2022-11-10 cs.LG cs.AI 62%

Understanding the Evolution of Linear Regions in Deep Reinforcement Learning

Setareh Cohan, Nam Hee Kim, David Rolnick, Michiel van de Panne

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2022 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.10312 2022-09-20 cs.RO cs.AI cs.LG 62%

Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks

Bohan Wu, Suraj Nair, Li Fei-Fei, Chelsea Finn

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments Equal advising and contribution for last two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04665 2022-09-14 cs.AI cs.LG 62%

Ask Before You Act: Generalising to Novel Environments by Asking Questions

Ross Murphy, Sergey Mosesov, Javier Leguina Peral, Thymo ter Doest

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07646 2022-07-18 cs.CV cs.LG 62%

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Rui Qian, Yeqing Li, Zheng Xu, Ming-Hsuan Yang, Serge Belongie, Yin Cui

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00433 2022-07-04 cs.CV cs.LG 62%

PROTOtypical Logic Tensor Networks (PROTO-LTN) for Zero Shot Learning

Simone Martone, Francesco Manigrasso, Lamberti Fabrizio, Lia Morra

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09827 2022-06-22 cs.LG cs.AI 62%

A Distributional Approach for Soft Clustering Comparison and Evaluation

Andrea Campagner, Davide Ciucci, Thierry Denœux

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments This is the extended version of article "A Distributional Approach for Soft Clustering Comparison and Evaluation", accepted at BELIEF 2022 (http://hebergement.universite-paris-saclay.fr/belief2022/). Please cite the proceedings version of the article

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03774 2022-05-10 cs.CV cs.AI 62%

RoViST:Learning Robust Metrics for Visual Storytelling

Eileen Wang, Caren Han, Josiah Poon

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00974 2022-04-14 cs.CV cs.AI cs.CL 62%

Consensus Graph Representation Learning for Better Grounded Image Captioning

Wenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao, Qiang Yu, Yueting Zhuang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 5 figures, AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03647 2022-04-08 cs.CV cs.AI 62%

Adapting CLIP For Phrase Localization Without Further Training

Jiahao Li, Greg Shakhnarovich, Raymond A. Yeh

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11426 2022-04-07 cs.CV cs.GR cs.LG 62%

Neural Fields in Visual Computing and Beyond

Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, Srinath Sridhar

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Equal advising: Vincent Sitzmann and Srinath Sridhar

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16682 2022-04-01 cs.CV cs.CL cs.LG 62%

To Find Waldo You Need Contextual Cues: Debiasing Who's Waldo

Yiran Luo, Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

Comments Accepted at ACL 2022 (Short Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13344 2022-03-28 cs.CL cs.AI cs.LG 62%

Linking Emergent and Natural Languages via Corpus Transfer

Shunyu Yao, Mo Yu, Yang Zhang, Karthik R Narasimhan, Joshua B. Tenenbaum, Chuang Gan

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments ICLR 2022 Spotlight. Github repo: https://github.com/ysymyth/ec-nl

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12686 2022-03-25 cs.LG cs.AI cs.NE stat.ML 62%

Possibility Before Utility: Learning And Using Hierarchical Affordances

Robby Costales, Shariq Iqbal, Fei Sha

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

Comments ICLR 2022 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10568 2022-03-22 cs.RO cs.AI cs.CV 62%

Accelerating Integrated Task and Motion Planning with Neural Feasibility Checking

Lei Xu, Tianyu Ren, Georgia Chalvatzaki, Jan Peters

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI

Comments 6 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.08937 2022-03-11 cs.LG cs.CV 62%

When, Why, and Which Pretrained GANs Are Useful?

Timofey Grigoryev, Andrey Voynov, Artem Babenko

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏