arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本中的混沌:揭示多模态检索器的模态偏好

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers

Yubo Sun, Chunyi Peng, Yukun Yan, Zhenghao Liu, Zhipeng Xu, Sen Mei, Linlin Xin, Zheni Zeng, Maosong Sun

arXiv 2610.11816首次发表:更新:

发表机构

University of the Chinese Academy of Sciences; Northeastern University; Tsinghua University; Nanjing University(中国科学院大学; 东北大学; 清华大学; 南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对混合模态检索器的模态偏好问题,提出Trident方法缓解偏差,在多基准上提升了CLIP和VLM架构的混合模态检索性能。

AI 中文摘要

密集检索器在文本和图像语料库上已取得显著进展,但这些能力是否能可靠扩展到包含文本、图像及文本-图像融合文档的混合语料库尚不清楚。本文系统研究了不同架构的检索器,发现其性能对模态组成高度敏感:当图像文档被语义对应的文本表示逐步替换时,检索性能呈明显的V形曲线,在单模态语料库上表现强劲,但在模态共存时大幅下降。特别地,无关文本比同等数量的无关图像造成的退化更严重,我们将此现象称为“文本中的混沌”。进一步分析揭示了模态偏好:文本表示会获得系统性更高的相似度分数,导致无关文本排名高于相关图像。为缓解此偏差,我们提出Trident,它将每个文档的文本、图像及文本-图像融合视图构造为同等正样本,通过多正样本视图InfoNCE联合优化相关性判别与正视图平衡。在视觉文档和自然图像基准上的实验表明,Trident可提升基于CLIP和VLM架构的多模态检索性能,降低对模态组成和文本干扰项的敏感性,并提高平均单模态检索性能。

英文摘要

Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear. In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition. As image documents are progressively replaced with semantically corresponding text representations, retrieval performance follows a pronounced V-shaped curve, remaining strong on single-modality corpora but degrading substantially when modalities coexist. In particular, irrelevant text causes more severe degradation than an equal number of irrelevant images, a phenomenon we term Chaos in the Text. Further analysis reveals modality preference, whereby text representations receive systematically higher similarity scores, allowing irrelevant text to outrank relevant images. To mitigate this bias, we introduce Trident, which constructs text, image, and fused text-image views of each document as co-equal positives and jointly optimizes relevance discrimination and positive-view balance through Multi-Positive View InfoNCE. Experiments across visual document and natural image benchmarks show that trident improves mixed-modality retrieval on both CLIP-based and VLM-based architectures, reduces sensitivity to modality composition and text distractors, and increases average single-modality retrieval performance.

CommentsCode: https://github.com/OpenBMB/Trident

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑