arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3153 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3153 篇

2201.05017 2022-02-15 cs.CL cs.AI cs.LG 62%

Towards Automated Error Analysis: Learning to Characterize Errors

Tong Gao, Shivang Singh, Raymond J. Mooney

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04306 2022-02-10 cs.AI cs.CL cs.CV 62%

Can Open Domain Question Answering Systems Answer Visual Knowledge Questions?

Jiawen Zhang, Abhijit Mishra, Avinesh P. V. S, Siddharth Patwardhan, Sachin Agarwal

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 9 pages (including references), 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.08170 2022-01-19 cs.LG cs.CV 62%

How Modular Should Neural Module Networks Be for Systematic Generalization?

Vanessa D'Amario, Tomotake Sasaki, Xavier Boix

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Journal ref 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Sydney, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03342 2022-01-11 cs.CV cs.LG cs.MM 62%

COIN: Counterfactual Image Generation for VQA Interpretation

Zeyd Boukhers, Timo Hartmann, Jan Jürjens

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05587 2021-12-16 cs.CV cs.CL cs.LG 62%

Unified Multimodal Pre-training and Prompt-based Tuning for Vision-Language Understanding and Generation

Tianyi Liu, Zuxuan Wu, Wenhan Xiong, Jingjing Chen, Yu-Gang Jiang

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11832 2021-12-16 cs.CV cs.CL cs.LG 62%

Playing Lottery Tickets with Vision and Language

Zhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen, Yu Cheng, Shuohang Wang, Jingjing Liu, Lijuan Wang, Zicheng Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted to AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06416 2021-11-02 cs.CV cs.LG 62%

MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants

Alkesh Patel, Joel Ruben Antony Moniz, Roman Nguyen, Nick Tzou, Hadas Kotek, Vincent Renkens

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Extended abstract accepted for WeCNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14712 2021-10-27 cs.CV cs.AI cs.CY cs.HC 62%

Generating and Evaluating Explanations of Attended and Error-Inducing Input Regions for VQA Models

Arijit Ray, Michael Cogswell, Xiao Lin, Kamran Alipour, Ajay Divakaran, Yi Yao, Giedrius Burachas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Applied AI Letters, Wiley, 25 October 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06863 2021-10-19 cs.CV cs.AI cs.HC 62%

Improving Users' Mental Model with Attention-directed Counterfactual Edits

Kamran Alipour, Arijit Ray, Xiao Lin, Michael Cogswell, Jurgen P. Schulze, Yi Yao, Giedrius T. Burachas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in Applied AI Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.04484 2021-09-20 cs.CV cs.CL cs.LG 62%

Are VQA Systems RAD? Measuring Robustness to Augmented Data with Focused Interventions

Daniel Rosenberg, Itai Gat, Amir Feder, Roi Reichart

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments ACL 2021. Our code and data are available at https://danrosenberg.github.io/rad-measure/

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00451 2021-08-13 cs.CV cs.CL cs.LG 62%

Just Ask: Learning to Answer Questions from Millions of Narrated Videos

Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, Cordelia Schmid

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted at ICCV 2021 (Oral); 20 pages; 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04938 2021-08-12 cs.CV cs.AI cs.CL 62%

BERTHop: An Effective Vision-and-Language Model for Chest X-ray Disease Diagnosis

Masoud Monajatipoor, Mozhdeh Rouhsedaghat, Liunian Harold Li, Aichi Chien, C. -C. Jay Kuo, Fabien Scalzo, Kai-Wei Chang

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 8 figures, Accepted in ICCV workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10446 2021-06-22 cs.CV cs.AI 62%

Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering

Ahjeong Seo, Gi-Cheon Kang, Joonhan Park, Byoung-Tak Zhang

专题命中 视觉问答 :grounding(abstract);分类 cs.CV、cs.AI

Comments ACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12852 2021-06-18 cs.CV cs.AI 62%

Beyond VQA: Generating Multi-word Answer and Rationale to Visual Questions

Radhika Dua, Sai Srinivas Kancheti, Vineeth N Balasubramanian

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments MULA Workshop, CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.04630 2021-06-10 cs.CV cs.CL cs.LG 62%

PAM: Understanding Product Images in Cross Product Category Attribute Extraction

Rongmei Lin, Xiang He, Jie Feng, Nasser Zalmout, Yan Liang, Li Xiong, Xin Luna Dong

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments KDD 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.06365 2021-04-14 cs.LG cs.CV 62%

Neuro-Symbolic VQA: A review from the perspective of AGI desiderata

Ian Berlot-Attwell

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03046 2021-04-08 cs.CV cs.LG 62%

Multimodal Continuous Visual Attention Mechanisms

António Farinhas, André F. T. Martins, Pedro M. Q. Aguiar

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02096 2021-04-07 cs.CV cs.AI 62%

Compressing Visual-linguistic Model via Knowledge Distillation

Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, Zicheng Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01394 2021-04-06 cs.CV cs.CL cs.LG 62%

MMBERT: Multimodal BERT Pretraining for Improved Medical VQA

Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, CV Jawahar

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11537 2021-03-23 cs.LG cs.CV 62%

How to Design Sample and Computationally Efficient VQA Models

Karan Samel, Zelin Zhao, Binghong Chen, Kuan Wang, Robin Luo, Le Song

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05900 2021-03-11 cs.CV cs.AI 62%

RL-CSDia: Representation Learning of Computer Science Diagrams

Shaowei Wang, LingLing Zhang, Xuan Luo, Yi Yang, Xin Hu, Jun Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.06793 2021-02-16 cs.CV cs.AI cs.CL 62%

Unanswerable Questions about Images and Texts

Ernest Davis

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 4 figures

Journal ref Frontiers in Artificial Intelligence: Language and Computation. July 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.00424 2021-02-02 cs.CL cs.CV cs.LG 62%

An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games

Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted paper for the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14891 2021-01-01 cs.CV cs.LG 62%

Detecting Hate Speech in Multi-modal Memes

Abhishek Das, Japsimar Singh Wahi, Siyao Li

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07214 2020-10-30 cs.LG cs.CL cs.CV stat.ML 62%

Sparse and Continuous Attention Mechanisms

André F. T. Martins, António Farinhas, Marcos Treviso, Vlad Niculae, Pedro M. Q. Aguiar, Mário A. T. Figueiredo

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted for spotlight presentation at NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10802 2020-10-22 cs.CV cs.LG 62%

Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies

Itai Gat, Idan Schwartz, Alexander Schwing, Tamir Hazan

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10354 2020-10-13 cs.CV cs.CL cs.LG 62%

Unsupervised Keyword Extraction for Full-sentence VQA

Kohei Uehara, Tatsuya Harada

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments EMNLP 2020 workshop: NLP Beyond Text (NLPBT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00562 2020-10-02 cs.CL cs.AI cs.CV 62%

ISAAQ -- Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down Attention

Jose Manuel Gomez-Perez, Raul Ortega

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication as a long paper in EMNLP2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.00124 2020-08-31 cs.CV cs.CL cs.HC cs.LG 62%

A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop

Jihyeon Lee, Sho Arora

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 9 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07194 2020-08-05 cs.CL cs.CV cs.IR cs.LG cs.MM 62%

Recommending Themes for Ad Creative Design via Visual-Linguistic Representations

Yichao Zhou, Shaunak Mishra, Manisha Verma, Narayan Bhamidipati, Wei Wang

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 7 pages, 8 figures, 2 tables, accepted by The Web Conference 2020

详情

展开后加载摘要…

URL PDF HTML 收藏