AI 中文总结
该研究针对多模态问答的模态偏差、知识不确定性与推理深度不足问题,提出多模态生成模糊系统(MMGFS),通过多模态协同沉思机制与模糊规则多跳推理实现性能提升,在多数据集上优于现有方法。
AI 中文摘要
在多模态问答(MQA)任务中,模型需联合编码并整合来自文本、图像、语音等多模态的异构信息,以完成复杂的语义推理与决策。尽管近期取得了诸多进展,现有方法(包括传统深度学习模型、大模型(LMs)或基于提示的框架)仍面临若干关键挑战:其一,模态偏差源于不同模态间的特征分布差异,限制了有效的跨模态协同理解;其二,许多问题需要融合多领域知识,带来显著的不确定性;其三,当前方法常依赖浅层语义匹配,导致推理深度有限、可解释性降低。为解决这些问题,受传统模糊系统(FS)框架启发,我们提出了一种模糊推理引导的多模态生成架构,命名为多模态生成模糊系统(MMGFS)。MMGFS的主要贡献有两点:一是通过多模态协同沉思机制缓解模态偏差;二是引入模糊规则与多跳推理机制,支持跨领域知识融合与分层推理,从而强化不确定性建模并深化语义理解。我们在开放域问答数据集(包括MultimodalQA和WebQA)以及特定领域基准(包括BioMol-VQA和EHRxQA)上开展了全面评估,实验结果表明,MMGFS在多个数据集上始终优于现有方法,可有效缓解模态偏差与问题不确定性,同时在答案准确率、一致性和泛化性上实现了更优性能。
英文摘要
In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision making. Despite recent advances, existing approaches, including traditional deep learning models and Large Models (LMs) or prompt-based frameworks, continue to face several critical challenges. First, modality bias arises from discrepancies in feature distributions across different modalities, which limits effective cross modal collaborative understanding. Second, many questions require knowledge drawn from multiple domains, introducing significant uncertainty. Third, current methods often rely on shallow semantic matching, resulting in limited reasoning depth an reduced interpretability. To address these issues, inspired by the traditional fuzzy system (FS) framework, we propose a fuzzy-inference-guided multimodal generative architecture termed the Multi-Modal Generative Fuzzy System (MMGFS). The main contributions of MMGFS are two folds. First, it alleviates modality bias through a multimodal collaborative rumination mechanism. Second, it introduces fuzzy rules and a multi-hop inference mechanism to support cross-domain knowledge fusion and hierarchical reasoning, thereby strengthening uncertainty modelling and deepening semantic understanding. We conduct comprehensive evaluations on open-domain question answering datasets, including MultimodalQA and WebQA, as well as domain-specific benchmarks, including BioMol-VQA and EHRxQA. Experimental results demonstrate that MMGFS consistently outperforms existing methods across multiple datasets. It effectively mitigates modality bias and question uncertainty while achieving superior performance in answer accuracy, consistency, and generalization.
Comments13 pages, 8 figures