AI 中文总结
研究多语言检索增强生成中重排器对文档语言的考量问题,提出语言感知多语言交叉编码器LAMAR,它通过英语锚定相关性蒸馏和偏好对齐实现语言连贯,在控制实验及实际检索中表现出色,证明其兼顾语言连贯与重排性能。
AI 中文摘要
在多语言检索增强生成中,检索器可检索多种语言的相关文档,在答案生成前进行重排。但现有多语言重排器在对语义相关候选文档排序时是否考虑文档语言尚不明晰。分析表明,即便文档语言会影响答案生成,当跨语言存在语义等效文档时,这些重排器并非始终优先考虑与查询语言相同的文档。我们发布了LAMAR,一种语言感知多语言交叉编码器,它通过英语锚定相关性蒸馏在多语言输入间建立一致的相关性评分,再应用偏好对齐以实现语言连贯,促使与查询语言相同的文档在保持语义相关性的同时获得更高排名。在评估语言连贯的控制实验中,LAMAR总体及各语言单独评估时均表现最佳,在既定多语言重排基准测试中也具竞争力,在实际检索设置中重排第一阶段检索出的候选文档时,LAMAR在所有报告指标上取得最佳结果。这些结果表明LAMAR在实现强大的多语言重排基准性能的同时,考虑了语言连贯。
英文摘要
In multilingual retrieval augmented generation pipelines, an embedding model can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document language when ordering semantically relevant candidates. Our analysis shows that these rerankers do not consistently prioritize documents written in the same language as the query when semantically equivalent documents are available across languages, even though document language can affect answer generation. We release LAMAR, a language aware multilingual cross encoder trained to account for both semantic relevance and language coherence. LAMAR first uses English anchored relevance distillation to establish consistent relevance scoring across multilingual inputs and then applies preference alignment for language coherence to encourage documents written in the same language as the query to receive higher rankings while retaining semantic relevance. In a controlled experiment designed to assess language coherence, LAMAR achieves the best performance overall and across all languages examined individually. LAMAR also remains competitive on established multilingual reranking benchmarks. In practical retrieval settings, LAMAR achieves the best results across all reported metrics when reranking candidates retrieved in the first stage. These results demonstrate that LAMAR accounts for language coherence while achieving strong performance on general multilingual reranking benchmarks.
Commentspreprint