发表机构
ConfidentialMind(ConfidentialMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出查询感知路由结合低秩适配器,在不损害同语言性能的前提下,显著提升芬兰语和瑞典语的跨语言检索质量,相对增益达20.9%。
AI 中文摘要
多语言编码器在查询和相关文档语言不同时,尽管同语言性能强劲,但检索效果可能降低。我们研究了芬兰语和瑞典语的跨语言检索是否能在保持编码器现有同语言性能和文档索引的同时得到提升。我们将一个仅针对查询的低秩适配器(针对冻结的文档嵌入进行训练)与基于查询和索引语言的确定性路由相结合。跨语言查询使用适配器,而同语言查询则使用原始编码器。我们的SampoTron,即微调的低秩(LoRA)适配器与Nemotron-3-Embed-1B模型配合,在六个英语、芬兰语和瑞典语方向上的平均检索质量从排名前十的归一化折损累积增益(nDCG)0.241提升至0.291,相对提升了20.9%,这是在采样的金融基准上实现的。所有六个跨语言方向均有所提升,且路由保持了原始的同语言性能,包括两项全语料库芬兰语评估。该方法实现了选择性的跨语言专业化,同时可复用文档嵌入向量。
英文摘要
Multilingual encoders can exhibit reduced retrieval effectiveness when queries and relevant documents differ in language, despite strong same-language performance. We investigate whether Finnish and Swedish cross-lingual retrieval can improve while preserving an encoder's existing same-language performance and document index. We combine a query-only low-rank adapter, trained against frozen document embeddings, with deterministic routing based on query and index languages. Cross-language queries use the adapter, while same-language queries use the original encoder. SampoTron, our fine-tuned low-rank (LoRA) adapter alongwith the Nemotron-3-Embed-1B model, improves average retrieval quality across six English, Finnish, and Swedish directions from 0.241 to 0.291 in normalized discounted cumulative gain (nDCG) at rank ten, a 20.9% relative gain on a sampled financial benchmark. All six cross-lingual directions improve, and routing preserves the original same-language performance, including two full-corpus Finnish evaluations. The approach enables selective cross-language specialization with reusable document embedding vectors.