arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26349 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3153 篇

2104.00926 2021-07-21 cs.CV cs.HC 57%

VisQA: X-raying Vision and Language Reasoning in Transformers

Theo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05556 2021-07-09 cs.CL cs.CV 57%

Sparse and Structured Visual Attention

Pedro Henrique Martins, Vlad Niculae, Zita Marinho, André Martins

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09141 2021-06-18 cs.CL cs.CV 57%

Probing Image-Language Transformers for Verb Understanding

Lisa Anne Hendricks, Aida Nematzadeh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05251 2021-06-10 cs.LG cs.CL stat.ML 57%

Bayesian Attention Belief Networks

Shujian Zhang, Xinjie Fan, Bo Chen, Mingyuan Zhou

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

Comments ICML 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07141 2021-05-18 cs.CV 57%

Show Why the Answer is Correct! Towards Explainable AI using Compositional Temporal Attention

Nihar Bendre, Kevin Desai, Peyman Najafirad

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 7 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.01993 2021-05-06 cs.CV 57%

AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss

Yangyang Guo, Liqiang Nie, Zhiyong Cheng, Feng Ji, Ji Zhang, Alberto Del Bimbo

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06087 2021-04-20 cs.CV 57%

Contrast and Classify: Training Robust VQA Models

Yash Kant, Abhinav Moudgil, Dhruv Batra, Devi Parikh, Harsh Agrawal

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08108 2021-04-19 cs.CV cs.CL 57%

Cross-Modal Retrieval Augmentation for Multi-Modal Classification

Shir Gur, Natalia Neverova, Chris Stauffer, Ser-Nam Lim, Douwe Kiela, Austin Reiter

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03762 2021-04-09 cs.CV cs.CL 57%

Video Question Answering with Phrases via Semantic Roles

Arka Sadhu, Kan Chen, Ram Nevatia

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments NAACL21 Camera Ready including appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03656 2021-04-09 cs.CV 57%

How Transferable are Reasoning Patterns in VQA?

Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche, Romain Vuillemot, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05726 2021-04-09 cs.CV cs.CL 57%

Estimating semantic structure for the VQA answer space

Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments [WARNING] We want to notice the reader that additional experiments (not in the paper) have shown that using a `random' semantic space performs as much as the proposed semantic loss. This additional result question the effectiveness of our method

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05121 2021-04-08 cs.CV 57%

Roses Are Red, Violets Are Blue... but Should Vqa Expect Them To?

Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.00332 2021-04-02 cs.CV 57%

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, Jingjing Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15974 2021-03-31 cs.CV 57%

Domain-robust VQA with diverse datasets and methods but no target labels

Mingda Zhang, Tristan Maidment, Ahmad Diab, Adriana Kovashka, Rebecca Hwa

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments To appear in CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.08981 2021-03-31 cs.CV cs.CL 57%

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Soravit Changpinyo, Piyush Sharma, Nan Ding, Radu Soricut

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2021). Our dataset is available at https://github.com/google-research-datasets/conceptual-12m

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04181 2021-03-09 cs.LG 57%

Contextual Dropout: An Efficient Sample-Dependent Dropout Module

Xinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, Mingyuan Zhou

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

Journal ref ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.06573 2021-01-19 cs.AI 57%

Understanding in Artificial Intelligence

Stefan Maetschke, David Martinez Iraola, Pieter Barnard, Elaheh ShafieiBavani, Peter Zhong, Ying Xu, Antonio Jimeno Yepes

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments 28 pages, 282 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.05153 2020-12-10 cs.CV 57%

Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps

Qi Zhu, Chenyu Gao, Peng Wang, Qi Wu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02951 2020-12-08 cs.CV 57%

FloodNet: A High Resolution Aerial Imagery Dataset for Post Flood Scene Understanding

Maryam Rahnemoonfar, Tashnim Chowdhury, Argho Sarkar, Debvrat Varshney, Masoud Yari, Robin Murphy

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04963 2020-12-02 cs.CV 57%

Rephrasing visual questions by specifying the entropy of the answer distribution

Kento Terao, Toru Tamaki, Bisser Raytchev, Kazufumi Kaneda, Shun'ichi Satoh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14759 2020-11-30 cs.AI 57%

Graph-based Heuristic Search for Module Selection Procedure in Neural Module Network

Yuxuan Wu, Hideki Nakayama

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments in Neural Module Network[C]//Proceedings of the Asian Conference on Computer Vision. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11721 2020-11-25 cs.CV 57%

Siamese Tracking with Lingual Object Constraints

Maximilian Filtenborg, Efstratios Gavves, Deepak Gupta

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.11894 2020-11-24 cs.CV 57%

Unshuffling Data for Improved Generalization

Damien Teney, Ehsan Abbasnejad, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04264 2020-11-10 cs.CL cs.CV 57%

CapWAP: Captioning with a Purpose

Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark, Regina Barzilay

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.13128 2020-10-27 cs.AI cs.CL cs.IR 57%

ExplanationLP: Abductive Reasoning for Explainable Science Question Answering

Mokanarangan Thayaparan, Marco Valentino, André Freitas

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11701 2020-10-23 cs.CV cs.CL 57%

Spatial Attention as an Interface for Image Captioning Models

Philipp Sadler

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments A thesis submitted in fulfillment of the requirements for the degree Master of Science in Cognitive Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10604 2020-10-22 stat.ML cs.LG cs.NE 57%

Bayesian Attention Modules

Xinjie Fan, Shujian Zhang, Bo Chen, Mingyuan Zhou

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.08189 2020-10-19 cs.CV 57%

New Ideas and Trends in Deep Multimodal Content Understanding: A Review

Wei Chen, Weiping Wang, Li Liu, Michael S. Lew

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06572 2020-10-14 cs.CL cs.CV 57%

Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!

Jack Hessel, Lillian Lee

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Journal ref Published in EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14451 2020-10-07 cs.CL cs.CV 57%

Pragmatic Issue-Sensitive Image Captioning

Allen Nie, Reuben Cohn-Gordon, Christopher Potts

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 15 pages, 7 figures. EMNLP 2020 Findings Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏