MERGE:通过生成式增强实现多LLM检索集成
MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
- University of Waterloo(滑铁卢大学)
- National Taiwan University(国立台湾大学)
- Rakuten Group, Inc.(乐天集团株式会社)
- Rakuten Asia Pte. Ltd.(乐天亚洲私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MERGE框架,利用三个开源LLM生成查询扩展并由大模型集成,结合任务导向的自动提示优化,在BEIR基准上显著提升BM25检索性能。
AI中文摘要:
大型语言模型(LLMs)越来越多地被用于信息检索(IR)中丰富用户查询,以便诸如BM25之类的标准检索器能够弥合与目标语料库之间的词汇差距。然而,任何单一的LLM都受限于其训练数据和架构偏见,且其增强行为依赖于手工设计的提示,这些提示必须针对每个新模型重新设计——这是一个昂贵且扩展性差的过程。我们提出了MERGE(通过生成式增强实现多LLM检索集成),这是一个两阶段框架:三个异构的7-8B开源LLM独立生成候选扩展,一个更大的LLM生成式地将它们综合成单一查询。为了使提示工程在集成中具有可扩展性,我们将一个基于任务的自动提示优化(APO)循环集成到两个阶段中。与使用LLM评估器评判候选的APO方法不同,我们的循环根据每个候选的下游检索性能对其进行评分,并在当前冠军提示与优化器提出的草稿之间进行小型锦标赛,一旦冠军连续两轮存活即终止;一个历史增强变体还额外将最近的锦标赛轨迹反馈给优化器。MERGE与检索器无关,仅执行一次BM25传递,无需排名融合、无监督文档扩展和重新索引。在五个BEIR基准(NQ、SciFact、FiQA、Touche-2020、DBPedia)上,MERGE将BM25 nDCG@10相对于原始查询提高了+2.1至+14.9个百分点,并且尽管仅使用紧凑的开源模型,仍能匹配或超越强大的基于LLM的查询扩展基线。消融实验证实,第二阶段集成优于任何单个第一阶段LLM,且基于任务的APO将大型种子提示回归转化为一致的增益,无需手动调整。
英文摘要:
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model -- an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce candidate expansions, and a larger LLM generatively synthesizes them into a single query. To make prompt engineering scalable across the ensemble, we integrate a task-grounded Automatic Prompt Optimization (APO) loop into both stages. Unlike APO methods that judge candidates with an LLM evaluator, our loop scores each candidate by its downstream retrieval performance and runs a small tournament between the current champion prompt and optimizer-proposed drafts, terminating once the champion survives two consecutive rounds; a history-augmented variant additionally feeds the recent tournament trajectory back to the optimizer. MERGE is retriever-agnostic and issues a single BM25 pass with no rank fusion, no supervised document expansion, and no re-indexing. On five BEIR benchmarks (NQ, SciFact, FiQA, Touche-2020, DBPedia), MERGE improves BM25 nDCG@10 over the original queries by +2.1 to +14.9 points and matches or outperforms strong LLM-based query-expansion baselines despite using only compact open-source models. Ablations confirm that the Stage-2 ensemble beats any single Stage-1 LLM, and that task-grounded APO converts large seed-prompt regressions into consistent gains without hand-tuning.