arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04949cs.CVcs.CLcs.ITcs.MMmath.IT

UG-UMRE:面向统一多模态关系抽取的不确定性引导模态增强与分布校准

UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

Bo Kong, Liruiz Jia, Yi Liang, Chao Liu, Dongfang Han, Tianwei Yan, Yuan Liu, Shengquan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对统一多模态关系抽取的噪声传播与模态分布异质性问题,提出含UDUA和JAUA模块的UG-UMRE网络,在UMRE等三个基准数据集上实现最优性能,模块可插拔且有效。

中文摘要 AI 辅助

统一多模态关系抽取(Unified Multimodal Relation Extraction, UMRE)旨在识别文本实体与视觉对象之间的模态内及跨模态关系。然而,现有UMRE研究仍面临两个关键问题:忽略固有偶然不确定性会导致噪声传播,而不同模态分布间的深层异质性则阻碍了对齐。为解决这些问题,我们提出了不确定性引导的UMRE网络(Uncertainty-Guided UMRE Network, UG-UMRE)。具体而言,我们设计了不确定性驱动的单模态增强(Uncertainty-Driven Unimodal Augmentation, UDUA)模块,该模块基于变分信息瓶颈将特征建模为高斯分布;通过引入不确定性感知的自监督对比学习机制,UDUA可有效过滤噪声同时保持语义一致性。此外,我们引入联合偶然不确定性对齐(Joint Aleatoric Uncertainty Alignment, JAUA)模块作为全局语义预校准机制,该模块利用概率分布一致性构建共享潜在空间,通过同步跨模态统计特性消除分布间隙,从而为细粒度交互奠定坚实基础。在三个基准数据集(UMRE、MORE和MNRE)上的实验表明,UG-UMRE实现了最优性能;进一步分析验证了所提出的UDUA和JAUA模块具有可插拔性且性能有效。

英文摘要

Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. However, existing UMRE studies still encounter two critical issues: ignoring inherent aleatoric uncertainty causes noise propagation, and deep-seated heterogeneity between distinct modal distributions hinders alignment. To address these issues, we propose the Uncertainty-Guided UMRE Network (UG-UMRE). Specifically, we design an Uncertainty-Driven Unimodal Augmentation (UDUA) module, which models features as Gaussian distributions based on the Variational Information Bottleneck. By incorporating an uncertainty-aware self-supervised contrastive learning mechanism, UDUA effectively filters out noise while maintaining semantic consistency. Furthermore, we introduce the Joint Aleatoric Uncertainty Alignment (JAUA) module as a global semantic pre-calibration mechanism. JAUA leverages probabilistic distribution consistency to construct a shared latent space, eliminating the distributional gap by synchronizing cross-modal statistical properties, thereby laying a robust foundation for fine-grained interaction. Experiments on three benchmark datasets (UMRE, MORE, and MNRE) demonstrate that UG-UMRE achieves state-of-the-art performance. Further analysis validates the pluggable and effective performance of the proposed UDUA and JAUA modules.

发表机构

  • Xinjiang University(新疆大学)
  • Chongqing Jiaotong University(重庆交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑