arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3153 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3153 篇

1909.11740 2020-07-21 cs.CV cs.CL cs.LG 62%

UNITER: UNiversal Image-TExt Representation Learning

Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, Jingjing Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.05608 2020-07-14 cs.CV cs.CL cs.LG 62%

Image Captioning with Compositional Neural Module Networks

Junjiao Tian, Jean Oh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments International Joint Conference on Artificial Intelligence (IJCAI-19)

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13073 2020-07-14 cs.CV cs.CL cs.LG 62%

A Novel Attention-based Aggregation Function to Combine Vision and Language

Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02509 2020-07-14 cs.LG cs.CV cs.NE 62%

REMIND Your Neural Network to Prevent Catastrophic Forgetting

Tyler L. Hayes, Kushal Kafle, Robik Shrestha, Manoj Acharya, Christopher Kanan

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments To appear in the European Conference on Computer Vision (ECCV-2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02833 2020-07-07 cs.CV cs.LG 62%

Eliminating Catastrophic Interference with Biased Competition

Amelia Elizabeth Pollard, Jonathan L. Shapiro

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.00900 2020-07-03 cs.CV cs.AI cs.HC 62%

The Impact of Explanations on AI Competency Prediction in VQA

Kamran Alipour, Arijit Ray, Xiao Lin, Jurgen P. Schulze, Yi Yao, Giedrius T. Burachas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Submitted to HCCAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.16917 2020-07-01 cs.AI cs.LG 62%

Ontology-guided Semantic Composition for Zero-Shot Learning

Jiaoyan Chen, Freddy Lecue, Yuxia Geng, Jeff Z. Pan, Huajun Chen

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

Comments Accepted by KR 2020 - 17th International Conference on Principles of Knowledge Representation and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.07538 2020-06-01 cs.CV cs.CL cs.LG 62%

Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing

Vedika Agarwal, Rakshith Shetty, Mario Fritz

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09241 2020-05-20 cs.CV cs.LG 62%

On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law

Damien Teney, Kushal Kafle, Robik Shrestha, Ehsan Abbasnejad, Christopher Kanan, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.09034 2020-04-21 cs.CV cs.LG 62%

Learning What Makes a Difference from Counterfactual Examples and Gradient Supervision

Damien Teney, Ehsan Abbasnedjad, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.08530 2020-02-19 cs.CV cs.CL cs.LG 62%

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, Jifeng Dai

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted by ICLR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03712 2020-01-14 cs.CV cs.CL cs.LG 62%

MHSAN: Multi-Head Self-Attention Network for Visual Semantic Embedding

Geondo Park, Chihye Han, Wonjun Yoon, Daeshik Kim

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted by the 2020 IEEE Winter Conference on Applications of Computer Vision (WACV 20), 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04964 2020-01-07 cs.MM cs.CV cs.IR cs.LG 62%

Multi-modal Deep Analysis for Multimedia

Wenwu Zhu, Xin Wang, Hongzhi Li

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 25 pages, 39 figures, IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03063 2019-12-09 cs.CV cs.CL cs.LG cs.NE 62%

Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks

Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.08618 2019-11-21 cs.CV cs.LG cs.MM eess.IV 62%

Explanation vs Attention: A Two-Player Game to Obtain Attention for VQA

Badri N. Patro, Anupriy, Vinay P. Namboodiri

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments AAAI-2020(Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.13471 2019-11-19 cs.CV cs.LG 62%

On Incorporating Semantic Prior Knowledge in Deep Learning Through Embedding-Space Constraints

Damien Teney, Ehsan Abbasnejad, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.06306 2019-10-18 cs.CV cs.CL cs.LG eess.IV 62%

U-CAM: Visual Explanation using Uncertainty based Class Activation Maps

Badri N. Patro, Mayank Lunayach, Shivansh Patel, Vinay P. Namboodiri

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments ICCV 2019 (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.04696 2019-09-12 cs.CV cs.AI 62%

Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation

Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, Giedrius Burachas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.03683 2019-09-10 cs.CL cs.CV cs.LG 62%

Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases

Christopher Clark, Mark Yatskar, Luke Zettlemoyer

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments In EMNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.04342 2019-08-16 cs.CV cs.HC cs.LG 62%

Why Does a Visual Question Have Different Answers?

Nilavra Bhattacharya, Qing Li, Danna Gurari

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Journal ref The IEEE International Conference on Computer Vision (ICCV) 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.00061 2019-08-02 cs.CV cs.LG 62%

An Empirical Study of Batch Normalization and Group Normalization in Conditional Computation

Vincent Michalski, Vikram Voleti, Samira Ebrahimi Kahou, Anthony Ortiz, Pascal Vincent, Chris Pal, Doina Precup

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.02314 2019-05-06 cs.LG cs.AI 62%

Answering Visual-Relational Queries in Web-Extracted Knowledge Graphs

Daniel Oñoro-Rubio, Mathias Niepert, Alberto García-Durán, Roberto González, Roberto J. López-Sastre

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Journal ref AKBC2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01385 2019-02-05 cs.LG cs.AI cs.CL cs.RO stat.ML 62%

Embodied Multimodal Multitask Learning

Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov, Devi Parikh, Dhruv Batra

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Comments See https://devendrachaplot.github.io/projects/EMML for demo videos

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.03928 2019-01-16 cs.LG cs.CV stat.ML 62%

Learning Representations of Sets through Optimized Permutations

Yan Zhang, Jonathon Hare, Adam Prügel-Bennett

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Published in ICLR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.12535 2018-11-01 cs.CV cs.AI 62%

Gated Hierarchical Attention for Image Captioning

Qingzhong Wang, Antoni B. Chan

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACCV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.12698 2018-10-31 cs.LG cs.AI cs.CL stat.ML 62%

Compositional Attention Networks for Interpretability in Natural Language Question Answering

Muru Selvakumar, Suriyadeepan Ramamoorthy, Vaidheeswaran Archana, Malaikannan Sankarasubbu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

Comments 8 pages,10 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.12366 2018-10-31 cs.AI cs.CL cs.CV 62%

Do Explanations make VQA Models more Predictable to a Human?

Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, Devi Parikh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2018. 16 pages, 11 figures. Content overlaps with "It Takes Two to Tango: Towards Theory of AI's Mind" (arXiv:1704.00717)

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.04344 2018-09-13 cs.CV cs.AI cs.CL 62%

The Wisdom of MaSSeS: Majority, Subjectivity, and Semantic Similarity in the Evaluation of VQA

Shailza Jolly, Sandro Pezzelle, Tassilo Klein, Andreas Dengel, Moin Nabi

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.09942 2018-08-30 cs.CL cs.AI cs.LG 62%

Neural Compositional Denotational Semantics for Question Answering

Nitish Gupta, Mike Lewis

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Comments EMNLP 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.04706 2018-05-22 cs.LG cs.CV cs.NE 62%

Tree Memory Networks for Modelling Long-term Temporal Dependencies

Tharindu Fernando, Simon Denman, Aaron McFadyen, Sridha Sridharan, Clinton Fookes

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Journal ref Neurocomputing, Volume 304, 23 August 2018, Pages 64-81

详情

展开后加载摘要…

URL PDF HTML 收藏