AI 中文总结
研究紧凑检索器能否从分数向量学教师排序行为,用行中心分数向量目标训练,在固定评估面板上蒸馏协议可弥补部分基础与教师模型差距,蒸馏后学生模型编码加速,外部迁移性能有好有坏。
AI 中文摘要
大型嵌入模型可提高检索质量,但在线运行大型编码器成本高昂。我们研究了一个紧凑的检索器能否在无法访问教师隐藏状态的情况下从分数向量中学习教师排序行为。学生模型在由我们的数据生成管道产生的真实正例和负候选构建的行上进行训练;我们将学生-教师硬负例挖掘作为扩展单独评估。我们使用以行为中心的分数向量目标,这是一种内存高效的均匀全对PairMSE损失实现。在固定的八项任务评估面板上,我们的蒸馏协议可弥补高达50%的基础模型与教师模型之间的差距。蒸馏后的0.6B学生模型在查询编码方面比顺序在线教师融合快4.7倍,在文档编码方面快9.7倍。蒸馏后的外部迁移性能好坏参半,因此我们的证据支持在匹配的检索协议下压缩教师排序。
英文摘要
Large embedding models improve retrieval quality, but serving large encoders online is expensive. We study whether a compact retriever can learn teacher ranking behavior from score vectors without access to teacher hidden states. The student trains on rows built from ground-truth positives and negative candidates produced by our data generation pipeline; we evaluate student-teacher hard-negative mining separately as an extension. We use a row-centered score-vector objective, a memory-efficient implementation of uniform all-pairs PairMSE loss. On a fixed eight-task evaluation panel, our distillation protocol recovers up to 50% of the base-to-teacher gap. The distilled 0.6B student is 4.7 times faster for query encoding and 9.7 times faster for document encoding than sequential online teacher fusion. External-transfer performance after distillation remains mixed, so our evidence supports compression of teacher rankings under matched retrieval protocols.