arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态问答的演进:从模态自适应提取到统一语言表示

Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation

Abdullah Al Shafi

arXiv 2609.08896首次发表:更新:

发表机构

Khulna University of Engineering & Technology(库尔纳工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文比较MAE、Solar和UniMMQA三种多模态问答框架,揭示从模态自适应处理向统一文本中心架构的演进,并证明该转变显著提升EM和F1分数,同时指出信息丢失等挑战。

AI 中文摘要

多模态数据的快速增长加剧了对能够跨异构来源(如文本、表格和图像)进行推理的问答(QA)系统的需求。在本文中,我们对三种有影响力的框架,即多模态自适应提取(MAE)、Solar和UniMMQA,进行了全面的方法论比较,追溯了多模态问答从模态自适应流水线到完全统一架构的演进过程。我们考察了每种方法如何建模跨模态交互、转换异构输入以及执行推理,重点突出了在模态表示、推理和答案生成方面的关键设计差异。我们的分析表明,从显式的模态特定处理向由预训练语言模型(PLMs)实现的统一文本中心公式化的明显转变。跨基准数据集的实证比较显示,这一转变在精确匹配(EM)和F1分数方面均带来了显著改进,其中UniMMQA实现了最一致且可扩展的性能。尽管取得了这些进展,我们仍识别出持续的挑战,包括模态转换过程中的信息丢失、多阶段流水线中的错误传播,以及捕捉细粒度跨模态依赖关系的局限性。总体而言,本研究提供了对当前设计趋势的更深入理解,并为统一多模态推理系统的未来方向提供了见解。

英文摘要

The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological comparison of three influential frameworks, namely Multimodal Adaptive Extraction (MAE), Solar, and UniMMQA, tracing the evolution of multimodal question answering from modality-adaptive pipelines to fully unified architectures. We examine how each approach models cross-modal interactions, transforms heterogeneous inputs, and performs reasoning, highlighting key design differences in modality representation, reasoning, and answer generation. Our analysis demonstrates a clear shift from explicit modality-specific processing toward unified text-centric formulations enabled by pre-trained language models (PLMs). Empirical comparisons across benchmark datasets show that this transition leads to substantial improvements in both Exact Match (EM) and F1-Scores, with UniMMQA achieving the most consistent and scalable performance. Despite these advances, we identify persistent challenges, including information loss during modality transformation, error propagation in multi-stage pipelines, and limitations in capturing fine-grained cross-modal dependencies. Overall, this study provides a deeper understanding of current design trends and offers insights into the future direction of unified multimodal reasoning systems.

Comments8 pages, 3 figures, 4 tables, reading assignment

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑