DDVT:用于视觉问答中答案定位的动态双级视觉Transformer融合网络
DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
浏览论文内容
中文总结 AI 辅助
研究视觉问答中答案定位问题,提出动态双级视觉Transformer融合网络DDVT,含问题引导的动态区域级模块和跨模态多尺度聚合模块,能精确定位相关视觉内容并融合文本特征,实验证明其性能优于现有方法。
中文摘要 AI 辅助
视觉问答中的答案定位旨在从与图像视觉内容相关的给定自然语言问题中定位区域,因其实际应用而备受关注。本文介绍了用于视觉问答中答案定位的动态双级视觉Transformer融合网络(DDVT)。具体而言,提出了问题引导的动态区域级模块(QGDR),通过ROI Align结合互补图像上下文和文本内容,实现与文本相关视觉内容的精确定位。还提出了跨模态多尺度聚合模块(CMA),增强像素级和区域级特征之间的融合,促进与有根据答案相关的视觉内容的有效定位。此外,将定位的视觉内容与文本特征融合以定位区域并回答关于图像的问题。实验结果表明,DDVT在几个广泛使用的基准上优于现有方法。
英文摘要
Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers. Furthermore, we fuse the located visual content with text features to locate the region and provide answers to questions posed about the image. Experimental results demonstrate that our DDVT outperforms state-of-the-art methods on several widely-used benchmarks.
发表机构
- National and Local Joint Engineering Laboratory of Computer Aided Design, School of Software Engineering, Dalian University(大连大学软件工程学院计算机辅助设计国家地方联合工程实验室)
- School of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。