arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02907cs.LG

贝叶斯数据重加权改进基于知识的视觉问答的多模态检索

Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering

  • University at Buffalo(布法罗大学)
  • NEC Laboratories America(美国 NEC 实验室)
  • Adobe Research(奥多比研究院)
  • Iowa State University(爱荷华州立大学)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Jingchen Sun, Shaobo Han, Ruiyi Zhang, Naresh Kumar Devulapally, Ming Liu, Yitao Long, Vishnu Suresh Lokhande, Changyou Chen

AI总结:

针对基于知识的视觉问答中多模态检索的负样本处理问题,提出贝叶斯数据重加权框架,通过概率建模与随机EM优化,在三个检索器和七个基准上提升了检索准确率。

AI中文摘要:

多模态检索器对基于知识的视觉问答至关重要,其为图像-问题对检索外部证据。但现有对比训练方法通常将所有不匹配的查询-文档对视为同等信息的负样本,这存在问题,因为许多不匹配文档仍可能在语义上相关或部分有用。我们提出贝叶斯数据重加权(Bayesian Data Reweighting),这是一个概率框架,将查询-文档重要性建模为潜在变量,并自适应推断后验权重以降低可能的假负样本权重。该方法在共轭先验下有闭式后验更新,并采用随机期望最大化(EM)优化,在三个检索器和七个基于知识的VQA基准上均一致提升了检索准确率。

英文摘要:

Multimodal retrievers are essential for knowledge-based visual question answering, where they retrieve external evidence for image-question pairs. However, existing contrastive training methods typically treat all unmatched query-document pairs as equally informative negatives, which is problematic because many unmatched documents may still be semantically relevant or partially useful. We propose Bayesian Data Reweighting, a probabilistic framework that models query-document importance as latent variables and adaptively infers posterior weights to downweight likely false negatives. With closed-form posterior updates under conjugate priors and stochastic EM optimization, our method consistently improves retrieval accuracy across three retrievers and seven knowledge-based VQA benchmarks.

↑