arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28968cs.AIcs.LG

面向语义搜索的高效GPU检索

Efficient GPU Retrieval for Semantic Search

Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bha… 展开作者

Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bhanuprasad Mehta, Meng Zheng, Puneet Singh Ahluwalia, Shirisha Singh, Ping Jin, Manas Apte, Gokulraj Mohanasundaram, Tugrul Bingol, Raghavan Muthuregunathan, Fedor Borisyuk

首次发表
浏览论文内容

中文总结 AI 辅助

针对领英语义搜索的瓶颈问题,提出策略对齐的GPU检索框架,经实验在多维度指标上显著提升搜索相关性与效率。

中文摘要 AI 辅助

领英平台的语义搜索需针对“柏林有支付领域工作经历的金融科技创始人”这类自然语言查询,从数亿规模的语料库中检索相关个人资料。已部署的相关性策略以瓶颈为导向:必须满足所有活跃且不可协商的维度,现有大语言模型(LLM)分级相关性(GR)评判器通过对维度评分进行固定的最小值/中位数聚合来实现该策略。而余弦相似度则对证据取平均值,会导致某一维度的强匹配掩盖另一维度的失败,从而限制第一阶段(L0)检索器的召回率。本文提出一种策略对齐的检索框架:将嵌入划分为8个类别监督段,其得分在服务时遵循相同的最小值/中位数规则;对于多向量检索,该段得分按每个带标签文档槽独立计算,并在各槽间取最大值。轻量级单槽第一阶段评分器生成高召回率候选,而尺度不变的相对范数门控则确保训练、评估和服务阶段的类别激活一致。在21000个保留查询上,该表示在匹配容量的基准之上提升了离线相关性,且增益广泛分布于各维度组合。我们采用两阶段GPU架构部署该框架:FP8粗排序器对全语料库评分,使每个分片的容量提升71%,第一阶段矩阵乘法吞吐量提升36%;随后FP16阶段对过采样候选集进行精确重排序,在每个分片副本超过500查询每秒(QPS)的情况下,恢复了全FP16召回率的99.6%-99.8%。在成员随机对照测试中,在不变的GR评判器下,探索性查询的Precision@10从63.7%提升至79.0%,导航性查询的Precision@1从65.5%提升至74.7%,盲法人工评估也独立证实了Precision@10的提升。

英文摘要

Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is bottleneck-oriented: every active non-negotiable facet must be satisfied, and a pre-existing LLM Graded Relevance (GR) judge operationalizes this through a fixed min/median aggregation over facet grades. Cosine similarity instead averages evidence, letting a strong match on one facet mask failure on another, capping the recall of the first-stage (L0) retriever. We present a policy-aligned retrieval framework: embeddings are partitioned into eight category-supervised segments whose scores follow the same min/median rule at serving time; for multi-vector retrieval, this segment score is computed independently per tagged document slot and maximized across slots. A lightweight single-slot Stage-1 scorer generates high-recall candidates, while scale-invariant relative-norm gating keeps category activation consistent across training, evaluation, and serving. On 21K held-out queries, this representation improves offline relevance over a matched-capacity baseline, with gains broadly distributed across facet combinations. We serve this framework with a two-stage GPU architecture: an FP8 coarse ranker scores the full corpus, increasing per-shard capacity by 71% and Stage-1 matmul throughput by 36%, then an FP16 stage exactly re-ranks an oversampled candidate set, recovering 99.6-99.8% of full-FP16 recall at over 500 QPS per shard replica. In a member-randomized A/B test, exploratory-query Precision@10 under the unchanged GR judge rises from 63.7% to 79.0% and navigational Precision@1 from 65.5% to 74.7%, with a blinded human evaluation independently confirming the Precision@10 gain.

发表机构

  • LinkedIn(领英)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑