Route Me If You Can: 查询改写选择基准
Route Me If You Can: A Benchmark for Query Reformulation Selection
浏览论文内容
中文总结 AI 辅助
该论文提出QueryRoute基准,用于评估查询改写选择问题,包含3,757个查询和619,905个检索结果,发现现有选择器仅部分恢复oracle提升空间,且排名随检索器变化。
中文摘要 AI 辅助
基于LLM的查询改写可以改善检索效果,但没有任何单一的改写策略在所有查询、领域、检索器或模型骨干上始终最优。这产生了一个推理时的决策问题:“给定原始查询和一组候选改写,应该将哪一个提交给检索器?”现有研究难以比较,因为它们使用不同的改写器池、检索器、相关性信号、训练标签和评估指标。我们引入了QueryRoute,一个冻结了可复现研究该决策所需昂贵工件的基准:原始查询、生成的变体、多个检索器下的排序列表、检索分数以及每个查询的oracle标签。该基准包含3,757个查询、11个候选系统、五个改写器骨干和三个检索器,覆盖TREC DL、BEIR和BRIGHT,产生619,905个检索结果。我们基准测试了监督分类、路由、QPP和LLM-as-judge选择器。结果显示,与固定改写器相比存在显著的oracle提升空间,但当前选择器仅恢复了其中一部分;选择器排名随检索器变化,相似的平均有效性可能掩盖不同的查询级行为。发布的工件和评估框架允许未来的选择器在不重新生成变体、不重新运行检索或重建评判管道的情况下进行比较。代码和数据可在https URL获取。
英文摘要
LLM-based query reformulation can improve retrieval, but no single reformulation strategy is consistently optimal across queries, domains, retrievers, or model backbones. This creates an inference-time decision problem: ``Given an original query and a pool of candidate reformulations, which one should be issued to the retriever?''. Existing studies are hard to compare because they use different reformulator pools, retrievers, relevance signals, training labels, and evaluation metrics. We introduce QueryRoute, a benchmark that freezes the expensive artifacts needed to study this decision reproducibly: original queries, generated variants, ranked lists under multiple retrievers, retrieval scores, and per-query oracle labels. The benchmark contains 3,757 queries, 11 candidate systems, five reformulator backbones, and three retrievers across TREC DL, BEIR, and BRIGHT, yielding 619,905 retrieval outcomes. We benchmark supervised classification, routing, QPP, and LLM-as-judge selectors. Results show substantial oracle headroom over fixed reformulators, but current selectors recover only part of it; selector rankings change across retrievers, and similar mean effectiveness can hide different query-level behavior. The released artifacts and evaluation harness allow future selectors to be compared without regenerating variants, rerunning retrieval, or rebuilding judge pipelines. Code and data are available at https://github.com/haisonle001/QueryRoute
发表机构
- Toronto Metropolitan University(多伦多都会大学)
- University of California, Berkeley(加州大学伯克利分校)
- University of Waterloo(滑铁卢大学)
- Mila - Quebec AI Institute(魁北克人工智能研究所)
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。