arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

参数高效的波兰语和欧洲语言检索器

Parameter-Efficient Retrievers for Polish and European Languages

Sławomir Dadas, Rafał Poświata, Małgorzata Grębowiec, Michał Perełkiewicz

arXiv 2609.12913首次发表:更新:

发表机构

National Information Processing Institute(国家信息处理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对稠密检索中大型模型的高成本问题,提出三阶段训练流程(跨语言对齐、知识蒸馏、对比微调)开发紧凑检索器PolDense和EuroDense,在41个波兰语和150个多语言任务上展现优异性能,PolDense-1B超越9B参数模型,EuroDense在低参数模型中领先。

AI 中文摘要

稠密检索系统日益依赖数十亿参数的语言模型,其内存和计算需求使得大规模索引、频繁的语料库更新以及低延迟服务成本高昂。我们提出了一种三阶段训练流程,用于开发紧凑且高效的检索器,这些检索器在与规模大得多的模型竞争时仍能保持竞争力。该流程结合了跨语言对齐、关系知识蒸馏和对比微调。它不需要原始的真实相关性标签,仅依赖于由强嵌入模型和重排序器(作为教师)生成的监督信号。利用这一流程,我们开发了PolDense和EuroDense,两者均支持长达8,192个标记的上下文。PolDense是一个包含六个波兰语检索器的系列,参数规模从17M到1B不等。EuroDense是一个支持九种欧洲语言的435M参数检索器。我们进行了广泛的评估,涵盖41个波兰语和150个多语言检索任务。结果表明了强大的质量-效率权衡。PolDense-1B在评估的检索器中优于参数高达9B的模型,而PolDense系列在模型规模上形成了帕累托前沿。在评估的低于1B参数的模型中,EuroDense在任务平均和语言平均性能上均排名第一,并在九种语言中的七种中领先。我们公开发布了所有模型。

英文摘要

Dense retrieval systems increasingly rely on multi-billion-parameter language models, whose memory and computational requirements make large-scale indexing, frequent corpus updates, and low-latency serving costly. We present a three-stage training pipeline for developing compact and efficient retrievers that remain competitive with substantially larger models. The pipeline combines cross-lingual alignment, relational knowledge distillation, and contrastive fine-tuning. It requires no original ground-truth relevance labels, relying exclusively on supervision generated by strong embedding models and rerankers utilised as teachers. Using this pipeline, we develop PolDense and EuroDense, both supporting contexts of up to 8,192 tokens. PolDense is a family of six Polish retrievers ranging from 17M to 1B parameters. EuroDense is a 435M-parameter retriever supporting nine European languages. We conduct an extensive evaluation covering 41 Polish and 150 multilingual retrieval tasks. The results demonstrate strong quality-efficiency trade-offs. PolDense-1B outperforms the evaluated retrievers with up to 9B parameters, while the PolDense family forms the Pareto frontier across model sizes. Among the evaluated models below 1B parameters, EuroDense ranks first in both task-averaged and language-averaged performance and leads in seven of nine languages. We release all models publicly.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑