arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39456cs.LGcs.CV

互均衡:通过互惠反馈的多模态表示学习

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

Ho-min Park, Byungkon Kang

首次发表
浏览论文内容

中文总结 AI 辅助

提出互反馈架构MEQ,通过两模态间持续信息交换形成耦合嵌入,以固定点作为输出,在分类与视觉接地任务上优于或媲美拼接方法。

中文摘要 AI 辅助

本工作提出了一种互反馈架构MEQ,该架构将两个可能不同模态的输入细化为一对耦合嵌入,使得每个嵌入都反映另一个嵌入的信息。核心思想是在两个输入之间持续交换信息。这一思想引出了一个由两个组件组成的互反馈架构,其中每个组件的输出被反馈到另一个组件。该模型的最终输出被定义为这种交互的固定点。我们提供了理论分析,对该模型进行了解释,并提供了防止失败情况的设计选择。我们通过跨多个数据集的分类和视觉接地任务展示了MEQ的优势。定量上,我们的模型在基于拼接的多模态分类问题上表现优于或与现有方法相当。定性上,所提出的交互机制允许模型在与互补模态配对时逐步细化视觉接地,从而展示了互反馈在这种设置下的强大能力。

英文摘要

This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the information of the other. The core idea is to incorporate continuous interchange of information between the two inputs. This idea leads to a mutual feedback architecture consisting of two components whose outputs are fed back into the other. The final output of this model is defined as the fixed point of this interaction. We provide theoretical analysis that offers interpretation of this model as well as design choices to prevent failure cases. We show the benefits of MEQ through classification and visual grounding tasks spanning various datasets. Quantitatively, our model outperforms or shows competitive performance on concatenation-based multimodal classification problems. Qualitatively, the proposed interactive mechanism allows the model to progressively refine the visual grounding when paired with complementary modality, thus demonstrating the power of mutual feedback under such settings.

发表机构

  • Baylor College of Medicine(贝勒医学院)
  • SUNY Korea(纽约州立大学韩国分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑