arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合专家语言模型可以成为强大且高效的检索器

Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers

Anubhav Shrestha, Safal Shrestha, Minwu Kim, Torsten Suel, Keith Ross

arXiv 2609.13486首次发表:更新:

发表机构

New York University Abu Dhabi; New York University(纽约大学阿布扎比分校; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文系统研究混合专家(MoE)语言模型作为检索器的性能,证明其以更少活跃参数和更低查询编码时间超越稠密检索器,并可通过减少专家数量进一步提速,成为高效的第一阶段检索器。

AI 中文摘要

近期研究表明,对仅解码器的大型语言模型(LLM)进行微调用于检索,可以产生强大的第一阶段检索器,且随着骨干模型规模的增大,有效性也随之提升。然而,每个查询和文档都必须通过整个模型,因此编码成本随模型规模增加而增加。混合专家(MoE)LLM 每个令牌仅激活参数的一个子集,被广泛用于扩展生成模型,但作为检索器仍未被充分探索。我们通过使用相同的训练流程,从多个模型家族训练 MoE 和稠密 LLM,在多样化的数据集上评估它们,并在相同的服务配置下测量查询编码时间,系统地研究了用于检索的 MoE 骨干模型。我们表明,MoE 检索器在 BEIR 上的 nDCG@10 得分比具有相当活跃参数数量的稠密检索器高出最多 3.0 个点。我们最强的 MoE 检索器之一以少 59% 的活跃参数和低 18% 的查询编码时间,匹配了 8B 稠密检索器的性能。我们进一步表明,无需重新训练或重新索引,即可减少用于查询编码的专家数量,在将查询编码时间最多降低 26% 的同时,保留了超过 99% 的检索有效性。最近的重新排序器在我们评估的配置中,相对于强大的 MoE 第一阶段仅提供适度的额外增益,且 MoE 第一阶段通常匹配或超过我们评估的重新排序配置。这些结果共同表明,MoE LLM 可以成为强大且高效的第一阶段检索器。

英文摘要

Recent work has shown that fine-tuning decoder-only large language models (LLMs) for retrieval yields strong first-stage retrievers, with effectiveness improving as backbones grow in size. However, every query and document must pass through the full model, so encoding cost increases with model size. Mixture-of-Experts (MoE) LLMs activate only a subset of parameters per token and are widely used to scale generative models, yet remain underexplored as retrievers. We systematically study MoE backbones for retrieval by training MoE and dense LLMs from several families using the same procedure, evaluating them across diverse datasets, and measuring query encoding time under the same serving configuration. We show that MoE retrievers outperform dense retrievers with comparable active parameter counts by up to 3.0 nDCG@10 points on BEIR. One of our strongest MoE retrievers matches an 8B dense retriever with 59% fewer active parameters and 18% lower query encoding time. We further show that the number of experts used for query encoding can be reduced without retraining or re-indexing, retaining more than 99% of retrieval effectiveness while reducing query encoding time by up to 26%. Recent rerankers provide only modest additional gains over strong MoE first stages, which often match or exceed the reranked configurations we evaluate. Together, these results show that MoE LLMs can be strong and efficient first-stage retrievers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑