arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用视觉语言模型推进Web规模搜索的相关性测量

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

Han Wang, Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath, Xi Chen, Roberto Konow, Kurchi Subhra Hazra

arXiv 2608.02446首次发表:更新:

发表机构

Pinterest(品趣志(Pinterest))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于VLM的自动化相关性评估流水线,部署于Pinterest搜索A/B实验,验证其与人工标注一致性,提升评估效率,降低MDEs,助力优化搜索体验。

AI 中文摘要

相关性评估在个性化搜索系统中发挥着关键作用,作为与用户参与指标并行的保障措施,确保搜索结果与用户查询及意图保持一致。人工标注是相关性评估的传统方法,但其高成本和长周转时间限制了可扩展性。本研究提出一种基于VLM(视觉语言模型)的自动化相关性评估流水线,已部署于Pinterest搜索中用于在线A/B实验。我们严格验证了VLM生成的判断与人工标注的一致性,证明VLMs可为实验提供可靠的相关性测量,同时大幅提升评估效率。利用基于VLM的标注还为扩展查询集、优化抽样设计及高效大规模评估更广泛的搜索体验创造了机会。该方法可得到更高质量的相关性指标,并显著降低在线实验测量中的最小可检测效应(MDEs)。

英文摘要

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.

CommentsRecSys'26 Industry track

DOI:10.1145/3773078.3831891

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑