arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ItColBERT:意大利语专用的后期交互检索器

ItColBERT: An Italian-Specialised Late-Interaction Retriever

Enrico Nello

arXiv 2609.26856首次发表:更新:

AI 中文总结

ItColBERT是首个意大利语专用后期交互检索器,以135M参数在四个基准上优于多数通用基线,并发现推理时分块在域外基准上带来显著提升,而额外训练和负样本策略未通过预注册门控。

AI 中文摘要

意大利语的神经信息检索几乎完全由多语言模型提供服务。多个多向量(后期交互)检索器将意大利语纳入其数十种语言之中,并且存在多个强大的意大利语稠密嵌入器,但截至2026年8月,尚未发布专门针对意大利语的后期交互检索器。我们提出了ItColBERT,这是一个135M参数的意大利语ColBERT模型,使用PyLate按照ColBERT-Zero配方进行训练:从已经能够检索的检查点初始化,然后进行监督对比训练,随后进行单教师蒸馏,总计在单个RTX 3090上约14.5 GPU小时。在四个意大利语检索基准上,它优于我们测试的所有通用后期交互基线,除了一个(mLateOn),并且参数数量比除一个规模相当的基线外的所有基线少2-4.4倍。我们的主要实证发现是方法论上的,且部分为负面。在唯一干净的域外基准(MLDR-it)上,对未改变的检查点应用推理时分块配方,nDCG@10提高了+0.0602(p = 0.0225),这一效果大于两轮进一步训练所产生的任何效果。自挖掘的困难负样本和原生1024 token训练均根据预注册的决策门进行评估,且均未通过。我们报告了所有与配对引导检验的比较,并针对经验测量的噪声下限0.0030 nDCG@10,同时我们发布了权重、训练和评估代码以及完整的实验记录,包括被拒绝的轮次。

英文摘要

Neural information retrieval for Italian is served almost entirely by multilingual models. Several multi-vector (late-interaction) retrievers include Italian among dozens of languages, and several strong Italian dense embedders exist, but as of August 2026 no late-interaction retriever specialised on Italian had been released. We present ItColBERT, a 135M-parameter Italian ColBERT trained with PyLate following the ColBERT-Zero recipe: initialise from a checkpoint that already retrieves, then apply supervised contrastive training followed by single-teacher distillation, for a total of roughly 14.5 GPU-hours on one RTX 3090. Across four Italian retrieval benchmarks it outperforms every general-purpose late-interaction baseline we tested except one (mLateOn), at 2-4.4x fewer parameters than every baseline but one of comparable size. Our principal empirical finding is methodological and partly negative. On the only cleanly out-of-domain benchmark (MLDR-it), an inference-time chunking recipe applied to an unchanged checkpoint yields +0.0602 nDCG@10 (p = 0.0225), a larger effect than anything two further rounds of training produced. Self-mined hard negatives and native 1024-token training were both evaluated against pre-registered decision gates and both failed. We report every comparison with paired bootstrap tests against an empirically measured noise floor of 0.0030 nDCG@10, and we release the weights, the training and evaluation code, and the complete experimental record including the rejected rounds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑