发表机构
University of Utah; University of Waterloo; The University of Queensland; Carnegie Mellon University; University of Virginia(犹他大学; 滑铁卢大学; 昆士兰大学; 卡内基梅隆大学; 弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Tevatron 3.0,集成Megatron-Core后端实现专家并行,在学术预算下成功训练30B参数MoE重排序器,其性能优于稠密模型且推理吞吐量更高。
AI 中文摘要
现代重排序方案——十亿参数级交叉编码器、混合专家(MoE)骨干网络以及针对强教师模型的知识蒸馏——已超出大多数学术团队可用的训练基础设施能力。现有的Tevatron重排序器训练依赖带有DeepSpeed或PyTorch FSDP1的Hugging Face Trainer,但这些后端对大规模MoE训练的高效支持不足。我们提出Tevatron 3.0,它将Megatron-Core训练后端集成到Tevatron中,同时保留其数据流水线、评估工作流以及与Hugging Face兼容的检查点。我们在新后端上对现有分布式训练配置进行基准测试,结果显示Megatron在可比的数据并行设置下,重排序器质量与FSDP相当,训练效率提升,在推荐的单节点配置下速度最高可达22%,且同时支持LoRA和全参数微调。关键的是,专家并行性使训练300亿参数的Qwen3-30B-A3B MoE重排序器成为可能,而这是PyTorch FSDP1无法实现的。利用该框架,我们在BEIR-15数据集上,针对三个第一阶段检索器,对MoE与稠密模型、LoRA与全参数微调、蒸馏与对比训练进行了受控对比,并报告了Hugging Face和vLLM的服务吞吐量。我们发现MoE重排序器与8B稠密模型质量相当,但激活的参数不到其一半,且推理吞吐量显著更高。我们将发布该框架和训练好的检查点。
英文摘要
Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced the training infrastructure available to most academic groups. Existing Tevatron reranker training relies on the Hugging Face Trainer with DeepSpeed or PyTorch FSDP1, but these backends lack efficient support for large-scale MoE training. We present Tevatron 3.0, which integrates a Megatron-Core training backend into Tevatron while preserving its data pipeline, evaluation workflow, and Hugging Face-compatible checkpoints. We benchmark existing distributed training configurations against the new backend, showing that Megatron matches FSDP reranker quality and training efficiency under comparable data-parallel settings, is up to 22% faster in the recommended single-node configuration, and supports both LoRA and full-parameter fine-tuning. Crucially, expert parallelism enables training a 30B-parameter Qwen3-30B-A3B MoE reranker, which is infeasible with PyTorch FSDP1. Using this framework, we conduct a controlled comparison of MoE versus dense models, LoRA versus full-parameter tuning, and distillation versus contrastive training on BEIR-15 with three first-stage retrievers, and report serving throughput for Hugging Face and vLLM. We find that the MoE reranker matches dense 8B quality while activating less than half as many parameters and achieving substantially higher inference throughput. We will release the framework and trained checkpoints.