发表机构
Walmart Global Technology; Walmart Global Tech(沃尔玛全球技术; 沃尔玛全球科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Doc2Query生成查询存在幻觉和重复问题,提出QGDPO方法,利用直接偏好优化和相关性过滤,显著减少不相关预测,已部署至生产环境并提升搜索相关性和用户参与度。
AI 中文摘要
Doc2Query是一种流行的文档扩展技术,利用序列到序列模型生成相关查询,有效解决了信息检索中的“词汇不匹配”问题。然而,这些模型常常生成与文档无关的幻觉内容或文档中已有的重复内容。训练序列到序列模型以产生高质量、新颖且相关的标记仍然是一个重大挑战。为解决这些问题,我们提出了一种新方法QGDPO,采用直接偏好优化(DPO)来指导生成过程。我们首先微调一个基础序列到序列模型,随后利用相关性模型对其预测进行评分。基于这些评分,我们构建胜出和失败预测对作为DPO训练的相关性偏好。此外,我们通过使用相关性模型过滤掉较差的预测来增强流程,仅保留最相关的生成内容用于索引。与Doc2Query基线相比,QGDPO有效消除了50%的不相关预测,而相关性过滤器额外移除了14.61%的不相关预测。该功能已成功部署到生产环境,在http URL上全流量运行,显著提升了相关性和用户参与度。
英文摘要
Doc2Query, a popular document expansion technique, leverages sequence-to-sequence models to generate relevant queries, effectively addressing the "vocabulary mismatch" problem in information retrieval. However, these models often suffer from generating either hallucinations unrelated to the document or repetitive content already present in the document. Training sequence-to-sequence models to produce high-quality, novel, and relevant tokens remains a significant challenge. To address these issues, we introduce a novel approach, QGDPO, that employs direct preference optimization (DPO) to guide the generation process. We first fine-tune a base sequence-to-sequence model and subsequently utilize a relevance model to score its predictions. Based on these scores, we construct pairs of winning and losing predictions as relevance preferences for the DPO training. Furthermore, we enhance our pipeline by using the relevance model to filter out poor predictions, retaining only the most relevant generated content for indexing. QGDPO effectively eliminates 50% of irrelevant predictions comparing against Doc2Query baselines, while the relevance filter removes an additional 14.61%. This feature has been successfully deployed to production on Walmart.com for full traffic, with a substantial improvement in relevance and user engagement.