发表机构
ZooWork Team(ZooWork团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对电商排序中偏好信号难以监督的问题,提出ZooWork-ShopRanker系列重排序器,利用LLM评审小组标注偏好对并对齐训练,8B模型蒸馏至4B和0.6B,显著超越开放基线,并发布ShopRank-Bench基准。
AI 中文摘要
为通用网页检索训练的开放重排序器在迁移到电商场景时表现不佳,因为在电商中,排序决策不仅依赖于主题相关性,还依赖于用户偏好、商品约束和商品间的比较适配度。这些偏好信号难以在大规模上进行监督:真实的搜索流量提供了真实的查询和候选商品,但没有干净的成对标签。我们提出了ZooWork-ShopRanker,一个与带标签的购物偏好对齐的电商重排序器系列(参数量为0.6B、4B和8B)。训练对由来自不同模型家族的推理大语言模型(LLM)组成的评审小组标注,该小组充当偏好预言机,并采用位置去偏的评判和一致性分层,重排序器在这些标签上进行训练。对齐后的8B旗舰模型随后作为蒸馏教师,用于高效的4B和0.6B模型,这些模型拟合其分数并在评判对上进行锐化。为了衡量进展,我们引入了ShopRank-Bench,一个包含约10,000个私有流量偏好对的污染受限基准,涵盖两种文本格式,并按每个标签得到多少评判家族的支持进行分层。ZooWork-ShopRanker-8B和-4B显著优于最强的开放重排序器基线,每个模型都显著优于其未对齐的基础模型,而ZooWork-ShopRanker-0.6B则优于其同规模模型;这些优势在两种格式中均保持,并扩展到常见的MTEB基准。我们发布模型和双格式的ShopRank-Bench,以促进进一步研究。
英文摘要
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are difficult to supervise at scale: real search traffic provides authentic queries and candidates but no clean pairwise labels. We present ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, and 8B) aligned to judge-labeled shopping preference. Training pairs are labeled by a panel of reasoning large language models (LLMs) from different families acting as a preference oracle, with position-debiased judgments and agreement tiers, and the rerankers are trained on these labels. The aligned 8B flagship then serves as a distillation teacher for the efficient 4B and 0.6B models, which are fit to its scores and sharpened on judged pairs. To measure progress, we introduce ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label. ZooWork-ShopRanker-8B and -4B significantly outperform the strongest open reranker baseline, every model significantly beats its own un-aligned base, and ZooWork-ShopRanker-0.6B beats its size peer; the gains hold in both formats and extend to common MTEB benchmarks. We release the models and the dual-format ShopRank-Bench to facilitate further research.
Commentsproject page: \url{https://serendipityoneinc.github.io/look-bench-page/shoprank-bench.html}