arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21281cs.IRcs.DCcs.LGcs.PF

超大规模个性化搜索的GPU-CPU混合检索

Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale

  • Meta Platforms, Inc.(元平台公司)

机构由 AI 辅助整理,请以论文原文为准。

Hao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Ming… 展开作者

Hao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Minglei Chen, Shuai Ding, Li Fan, Xiao Kong

AI总结:

针对万亿级个性化搜索的规模悖论,提出GPU-CPU混合协同检索系统,GPU处理深度交互预排序,CPU覆盖更广库存,经A/B测试验证提升相关性与参与度。

AI中文摘要:

在万亿文档规模上对用户生成内容进行基于嵌入的检索,暴露了两种生产需求之间的尖锐冲突:一是对具有丰富用户意图的查询进行深度、表达性的个性化,二是在固定延迟和资源预算下对海量库存的广泛覆盖。我们将此称为个性化-规模悖论:将全部服务库存驻留在GPU内存中资源消耗过大,而CPU计算无法在延迟关键路径上执行相同的交互密集型模型。我们提出了一种混合GPU-CPU协同服务系统,通过编排而非新模型类别来解决这一悖论。高深度GPU路径在约十亿文档的精选在线池上融合了检索和交互预排序,而高广度CPU路径则使用轻量级个性化评分搜索独立选择的、大约大二十倍的在线库存。每个请求可以运行任一路径或两者;候选在共享下游排序前进行去重。该系统已投入生产。与旧版纯CPU配置的全系统A/B测试提高了模型评分的相关性和实质性参与度,而单独路径实验在其各自的部署范围内显示出积极价值。检索日志显示这些路径贡献了结构上不同的候选,生产服务测量表征了它们的延迟,匹配容量计划量化了将建模深度分配给GPU和库存广度分配给CPU的经济理由。总之,这些结果验证了一种实用的、可独立演进的深度-广度架构,用于超大规模个性化搜索。

英文摘要:

Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization-scale paradox: hosting the full serving inventory in GPU memory is too resource intensive, while CPU compute cannot execute the same interaction-heavy model on the latency-critical path. We present a hybrid GPU-CPU co-serving system that resolves the paradox through orchestration rather than a new model class. A high-depth GPU pathway fuses retrieval and interaction pre-ranking over a curated online pool on the order of a billion documents, while a high-breadth CPU pathway searches an independently selected online inventory roughly twenty times larger with lightweight personalized scoring. Either or both pathways can run per request; candidates are deduplicated before shared downstream ranking. The system is deployed in production. A full-system A/B test against the legacy CPU-only configuration improves model-scored relevance and substantive engagement, while separate pathway experiments show positive value at their own deployment scopes. Retrieval logs show that the pathways contribute structurally distinct candidates, production serving measurements characterize their latency, and a matched capacity plan quantifies the economic rationale for assigning modeling depth to GPUs and inventory breadth to CPUs. Together, these results validate a practical, independently evolvable depth-breadth architecture for ultra-large-scale personalized search.

补充信息

↑