发表机构
LightOn(莱特昂)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出开放端到端方案训练检索模型,构建DenseOn、LateOn等模型,在BEIR上创同规模最优,发现晚交互模型泛化性更优,公开了模型、数据与代码。
AI 中文摘要
当前最先进的检索模型越来越依赖封闭的训练数据,导致可复现性差距。本文提出一种用于训练检索模型的开放端到端方案,并研究通过翻译训练将英文监督信号迁移至多语言检索的效果。首先从34个公开来源的14亿对数据中重构并整理出6.65亿个英文对比预训练对,同时利用挖掘的难负例构建188万个监督微调对,训练得到两个1.49亿参数的模型:单向量密集模型DenseOn和ColBERT风格的晚交互模型LateOn。二者在BEIR数据集上的平均nDCG@10分别达到56.20和57.22,在该参数规模下创下新的最先进结果。随后将已验证的英文数据翻译成8种语言,得到包含跨语言样本的28亿对数据,并基于mmBERT-base构建两个3.07亿参数的模型mDenseOn和mLateOn。尽管共享主干网络、数据和目标,但二者的表示行为存在差异:密集模型在英文和翻译语言上表现强劲,但在翻译训练支持范围外性能下降;而晚交互模型对未见过的语言和文字表现出更好的泛化能力。这表明词元级匹配将翻译训练从目标语言扩展策略转变为多语言泛化方案。本文公开发布了相关模型、数据集和训练代码。
英文摘要
State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scripts. This suggests that token-level matching turns translate-train from a target-language expansion strategy into a multilingual generalization recipe. We publicly release the models, datasets, and training code.
Comments21 pages, 3 figures, 12 tables