发表机构
University of Science, Ho Chi Minh City; Vietnam National University, Ho Chi Minh City(胡志明市理科大学; 越南国立大学胡志明市分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种多层融合Transformer模型,通过交叉注意力模块聚合图像和文本的多模态特征,在ViVQA数据集上取得了优于基线的结果,推动了越南语视觉问答的发展。
AI 中文摘要
近几十年来,人工智能在理解图像和与图像交互方面取得了显著进展。该技术的重要应用之一是视觉问答(VQA),这是一个要求计算机以自然方式理解并回答关于图像的问题的研究领域。尽管针对英语的VQA已有广泛的研究和开发,但针对其他语言(尤其是越南语)的类似工作却很少。这一差距为越南语语境下VQA技术的进步带来了重大挑战和机遇。通过弥合这一差距,越南语VQA领域不仅丰富了人工智能研究的多样性,还能在教育、医疗保健和娱乐等多个领域实现实际应用,服务于全球越南语使用者。因此,探索和开发越南语VQA系统在推进计算机视觉与自然语言处理交叉领域的研究和实际应用方面具有巨大潜力。在本文中,我们提出了一种多层融合Transformer模型,利用交叉注意力模块将来自不同层的图像和文本的多模态特征组合成聚合表示。我们的架构使我们能够从低级到高级提取信息。通过详细的实验和消融研究,我们的模型在ViVQA数据集上针对越南语取得了优于竞争基线的有前景的结果。
英文摘要
In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.