arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Tevatron-Elastic:用于训练弹性检索器和重排序器的统一抽象

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

Yu Wang, Shengyao Zhuang, Xueguang Ma, Zongyu Wu, Jimmy Lin, Vivek Srikumar, Zhichao Xu

arXiv 2608.08809首次发表:更新:

发表机构

University of Utah; The University of Queensland; University of Waterloo; Pennsylvania State University(犹他大学; 昆士兰大学; 滑铁卢大学; 宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Tevatron-Elastic提出统一抽象,可训练支持多种规模的检索器与重排序器,验证显示其成本低且加速效果显著,发布相关框架与检查点。

AI 中文摘要

单一模型规模会给生产环境中的检索系统灵活性带来挑战:部分场景需要更快的速度,部分场景需要更小的索引,而合适的权衡会随工作负载变化。在信息检索(IR)语境下,基于Transformer的模型可通过三种方式缩小规模:减少层数、减少上层处理的token数量、生成更短的嵌入,每种方式节省不同的计算资源。这些选项此前被单独研究,各有对应的方法、代码和训练设置,难以组合或适配新模型。我们提出~\textsc{ours}(即Tevatron-Elastic),将三者纳入统一抽象:单个对象指定模型可运行的任意规模,简短的调度表列出训练时使用的规模。训练后生成一个可服务所有规模的检查点,部署时用户可选择任意规模。该抽象同时覆盖检索器和重排序器,以及编码器和解码器模型,因它通过Hugging Face Transformers已暴露的接口工作;新增骨干模型仅需配置更改,无需新的建模代码。此前的方法——Matryoshka嵌入、提前退出、2D Matryoshka(如Starbucks)及分层token压缩——均成为我们统一抽象的特例。该接口还支持Matryoshka LTC(MLTC),可在一个检索器检查点中联合训练多种token压缩比例。为验证我们的框架,我们在三个骨干模型和两个任务上训练了20个检查点:质量曲线平滑,一个检查点的成本仅比单规模训练的模型略高,对照研究确认了 wallclock 加速比。我们发布该框架和所有检查点,作为构建弹性检索系统的资源。

英文摘要

A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with the workload. In the context of information retrieval (IR), a transformer-based model can be made smaller in three ways---using fewer layers, passing fewer tokens through the upper layers, or producing a shorter embedding---and each way saves a different compute resource. These options have been studied one at a time, each as its own method with its own code and training setup, which makes them hard to combine or adapt to a new model. We present~\ours to bring all three under one simple abstraction: a single object names any size the model can run at, and a short schedule lists the sizes to train. Training then produces one checkpoint that serves all of those sizes, and at deployment the user picks any of them. The same abstraction covers both retrievers and rerankers and both encoder and decoder models, as it works through interfaces that Hugging Face transformers already expose; a new backbone is a configuration change, not new modeling code. Prior methods---Matryoshka embeddings, early exit, 2D~Matryoshka (e.g., Starbucks), and layerwise token compression---become special cases of our unified abstraction. The same interface also enables Matryoshka~LTC (MLTC), which jointly trains several token-compression ratios in one retriever checkpoint. To validate our framework, we train 20 checkpoints across three backbones and two tasks: the quality curves are smooth, one checkpoint costs little over a model trained for a single size, and a controlled study confirms the wallclock speedups. We release the framework and all checkpoints as a resource for building elastic retrieval systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑