ROCS:面向请求的计算共享,用于高效大规模推荐
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
浏览论文内容
中文总结 AI 辅助
本研究提出面向请求的计算共享(ROCS)范式,通过GLM、DCA、IKBO等技术优化推荐模型推理,在公开基准和生产 workload 中均实现效率提升且保持或改善预测质量,已大规模部署应用。
中文摘要 AI 辅助
现代推荐模型通过扩展特征交互模块和序列模块来提升预测质量,但生产环境的成本限制了系统的扩展程度。本研究提出面向请求的计算共享(Request-Oriented Compute Sharing, ROCS),这是一种建模与推理范式,它利用了推荐推理的独特属性:每个用户请求会对应大量候选项进行评估,而请求侧的特征可在所有候选项间共享。ROCS尽可能推迟请求-候选项交互的执行,隔离依赖于候选项的表示,且每个请求仅对模型的大部分部分评估一次,而非每个候选项评估一次,从而在保持或提升预测质量的同时大幅提高推理效率。为实现该范式,我们开发了广义层掩码(Generalized Layer Masking, GLM),用于在特征交互架构中强制实现候选项隔离;还开发了深度交叉注意力(Deep Cross Attention, DCA),将面向请求的共享扩展到序列架构。为支持高效的GPU部署,我们协同设计了内核内广播优化(In-Kernel Broadcast Optimization, IKBO),可显著加速ROCS模型的执行。在公开基准上的实验表明,ROCS在各类推荐主干模型中均能改善质量-效率的权衡。在生产规模的工作负载中,ROCS在检索模型上实现了最高3倍的每秒查询数(QPS)提升,且无质量下降;在短视频排序模型上,实现了50%的QPS提升,同时相对对数损失(LogLoss)改善0.5%。ROCS已部署在覆盖广告、自然流量场景,检索、排序阶段,且推理复杂度跨度超过两个数量级的大规模推荐系统中,在降低基础设施成本的同时取得了显著的线上收益。
英文摘要
Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.
发表机构
- Meta AI
机构由 AI 辅助整理,请以论文原文为准。