发表机构
University of Massachusetts Amherst; Databricks Mosaic Research(马萨诸塞大学阿默斯特分校; Databricks Mosaic研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AutoIndex框架通过搜索对文档进行多种操作的程序来优化文档表示,在CRUMB基准上评估,相比BM25基线显著提升召回率,证明文档表示应作为优化目标,而非固定预处理选择。
AI 中文摘要
我们提出了AutoIndex,这是一个用于学习表示程序的框架:将原始文档映射到检索系统所使用表示的可执行转换。AutoIndex不是调整检索器、重排器或少量预处理超参数,而是在索引前搜索对文档进行切片、丰富、归一化、重新加权或重组的程序。每次迭代时,它执行验证引导的程序搜索,智能体诊断当前程序的失败并合成候选更新,仅保留能提高检索质量的更新。我们在CRUMB(异构检索任务基准)上评估AutoIndex,所有实验中BM25固定不变。学习到的程序在所有8个任务上比静态全文档BM25基线提高了召回率,Recall@100平均提高8.4%,nDCG@10平均提高8.3%,最大增益分别为Recall@100的30.5%和nDCG@10的43.6%。这些结果表明文档表示不应被视为检索开始前固定的预处理选择,而应作为明确的优化目标。可通过此https URL获取重现我们结果的代码。
英文摘要
We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the representations exposed to a retrieval system. Rather than tuning retrievers, rerankers, or a small set of preprocessing hyperparameters, AutoIndex searches over programs that slice, enrich, normalize, reweight, or reorganize documents before indexing. At each iteration, AutoIndex performs validation-guided program search, in which agents diagnose failures of the current program and synthesize candidate updates, retaining only updates that improve retrieval quality under the resulting index. We evaluate AutoIndex on CRUMB, a benchmark of heterogeneous retrieval tasks, with BM25 held fixed across all experiments. The learned programs improve recall over a static full-document BM25 baseline on all 8 tasks, with average gains of +8.4% in Recall@100 and +8.3% in nDCG@10, and largest gains of +30.5% in Recall@100 and +43.6% in nDCG@10. These results suggest that document representation should not be treated as a fixed preprocessing choice made before retrieval begins, but as an explicit optimization target. Code to reproduce our results is available at https://github.com/auto-index/autoindex.