arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02189cs.IRcs.CL

用于零样本多语言密集检索的解耦对比学习

Disentangled Contrastive Learning for Zero-Shot Multilingual Dense Retrieval

Chao Huang, Yufeng Chen, Changhao Guan, Guang Yang, Dongze Chen, Kaiyu Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对低资源语言检索数据稀缺问题,提出解耦对比学习方法,将多语言表示分离为语义与语言子空间,在mMARCO、MIRACL上的零样本多语言密集检索性能优于多个强基线。

中文摘要 AI 辅助

多语言密集检索旨在基于统一检索模型处理不同语言的查询与文档,其挑战在于实现对低资源语言的鲁棒检索迁移,这类语言通常标注检索数据稀缺。尽管过往研究在多语言语义表示学习中将高资源监督迁移至低资源语言,但共享表示常纠缠语义与语言特征,可能干扰检索语义相关性优化。与现有在该纠缠下学习语言无关语义特征的方法不同,本文提出用于多语言密集检索的解耦对比学习(Disentangled Contrastive Learning, DCL)方法,将多语言表示分离为语义与语言子空间。具体而言,设计基于分层语义对齐和语言去偏对比学习的解耦优化目标,通过在句子和词元层面对齐跨语言的检索相关语义,同时在语言子空间捕捉语言特定差异,减少语言诱导的语义匹配干扰;将这些目标与检索目标联合优化,以促进从英文监督到多语言密集检索的稳定零样本迁移。在mMARCO和MIRACL上的大量实验表明,本文方法始终优于多个强基线,证明其有效性与泛化能力。

英文摘要

Multilingual dense retrieval aims to handle queries and documents across different languages based on a unified retriever model. The challenge lies in enabling robust retrieval transfer to low-resource languages where annotated retrieval data is often scarce. Although previous studies transfer high-resource supervision to low-resource languages in multilingual semantic representation learning, the shared representation often entangles semantic and linguistic features, which may interfere with optimizing semantic relevance for retrieval. Different from existing methods that focus on learning language-agnostic semantic features under such entanglement, we propose a disentangled contrastive learning~(DCL) method for multilingual dense retrieval by separating multilingual representations into semantic and linguistic subspaces. Specifically, we design disentangled optimization objectives based on hierarchical semantic alignment and language debiasing contrastive learning. By aligning retrieval-relevant semantics across languages at both sentence and token levels while capturing language-specific variations in the linguistic subspace, these objectives reduce language-induced interference in semantic matching. We jointly optimize them with the retrieval objective to facilitate stable zero-shot transfer from English supervision to multilingual dense retrieval. Extensive experiments on mMARCO and MIRACL show that our method consistently outperforms several strong baselines, demonstrating its effectiveness and generalization ability.

发表机构

  • Beijing Jiaotong University(北京交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑