面向语义搜索的高效GPU检索
Efficient GPU Retrieval for Semantic Search
浏览论文内容
中文总结 AI 辅助
针对领英语义搜索的瓶颈问题,提出策略对齐的GPU检索框架,经实验在多维度指标上显著提升搜索相关性与效率。
中文摘要 AI 辅助
领英平台的语义搜索需针对“柏林有支付领域工作经历的金融科技创始人”这类自然语言查询,从数亿规模的语料库中检索相关个人资料。已部署的相关性策略以瓶颈为导向:必须满足所有活跃且不可协商的维度,现有大语言模型(LLM)分级相关性(GR)评判器通过对维度评分进行固定的最小值/中位数聚合来实现该策略。而余弦相似度则对证据取平均值,会导致某一维度的强匹配掩盖另一维度的失败,从而限制第一阶段(L0)检索器的召回率。本文提出一种策略对齐的检索框架:将嵌入划分为8个类别监督段,其得分在服务时遵循相同的最小值/中位数规则;对于多向量检索,该段得分按每个带标签文档槽独立计算,并在各槽间取最大值。轻量级单槽第一阶段评分器生成高召回率候选,而尺度不变的相对范数门控则确保训练、评估和服务阶段的类别激活一致。在21000个保留查询上,该表示在匹配容量的基准之上提升了离线相关性,且增益广泛分布于各维度组合。我们采用两阶段GPU架构部署该框架:FP8粗排序器对全语料库评分,使每个分片的容量提升71%,第一阶段矩阵乘法吞吐量提升36%;随后FP16阶段对过采样候选集进行精确重排序,在每个分片副本超过500查询每秒(QPS)的情况下,恢复了全FP16召回率的99.6%-99.8%。在成员随机对照测试中,在不变的GR评判器下,探索性查询的Precision@10从63.7%提升至79.0%,导航性查询的Precision@1从65.5%提升至74.7%,盲法人工评估也独立证实了Precision@10的提升。
英文摘要
Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is bottleneck-oriented: every active non-negotiable facet must be satisfied, and a pre-existing LLM Graded Relevance (GR) judge operationalizes this through a fixed min/median aggregation over facet grades. Cosine similarity instead averages evidence, letting a strong match on one facet mask failure on another, capping the recall of the first-stage (L0) retriever. We present a policy-aligned retrieval framework: embeddings are partitioned into eight category-supervised segments whose scores follow the same min/median rule at serving time; for multi-vector retrieval, this segment score is computed independently per tagged document slot and maximized across slots. A lightweight single-slot Stage-1 scorer generates high-recall candidates, while scale-invariant relative-norm gating keeps category activation consistent across training, evaluation, and serving. On 21K held-out queries, this representation improves offline relevance over a matched-capacity baseline, with gains broadly distributed across facet combinations. We serve this framework with a two-stage GPU architecture: an FP8 coarse ranker scores the full corpus, increasing per-shard capacity by 71% and Stage-1 matmul throughput by 36%, then an FP16 stage exactly re-ranks an oversampled candidate set, recovering 99.6-99.8% of full-FP16 recall at over 500 QPS per shard replica. In a member-randomized A/B test, exploratory-query Precision@10 under the unchanged GR judge rises from 63.7% to 79.0% and navigational Precision@1 from 65.5% to 74.7%, with a blinded human evaluation independently confirming the Precision@10 gain.
发表机构
- LinkedIn(领英)
机构由 AI 辅助整理,请以论文原文为准。