发表机构
Japan Advanced Institute of Science and Technology; University of Science, Ho Chi Minh City; Vietnam National University, Ho Chi Minh City; Zalo Research Center(日本先端科学技术大学院大学; 胡志明市理科大学; 越南国家大学胡志明市分校; Zalo研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉问答中的可回答性预测,提出VT-Transformer模型,利用Transformer架构融合视觉与文本特征,在VizWiz 2020数据集上验证了其有效性与鲁棒性。
AI 中文摘要
视觉问答中的可回答性是一个新颖且有吸引力的任务,旨在预测多模态数据中图像与问题之间的可回答性分数。现有工作通常利用从视觉问答系统到可回答性的二元映射,这并未反映该问题的本质。结合我们将可回答性视为回归任务的考虑,我们提出了VT-Transformer,它通过Transformer架构利用视觉和文本特征。在VizWiz 2020数据集上的实验结果表明,与竞争性基线相比,VT-Transformer在视觉问答中的可回答性预测上具有有效性和鲁棒性。
英文摘要
Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.
DOI:10.1109/ICIP42928.2021.9506796