arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00938cs.IR

GRACE:面向实时广告检索的生成式推荐加速引擎

GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval

Zhou Fang, Yuhang Huang, Ang Zhang, Yihan He, Ruichao Xiao, Chao Li, Yavuz Yetim, Sibyl Yang, Xiaohan Wei, Fei Tian, Liang Wang, Chonglin Sun, Liyuan Li, Nathan Yan, Gaoxiang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出GRACE系统,通过生成式目标匹配解决广告检索合规性问题,优化解码器设计降低延迟,满足实时广告检索的合规、延迟及成本要求。

中文摘要 AI 辅助

将生成式推荐系统应用于高吞吐量实时广告检索场景时,会面临两大服务挑战:一是合规性,需确保生成的每条广告都符合广告主受众定向规则的请求要求;二是计算资源,需满足严格的延迟和GPU成本要求,同时能在每个请求中通过宽波束解码生成数千条广告。本文提出GRACE,一种用于广告生成式检索的服务系统,可同时解决上述两个挑战。针对合规性,GRACE引入生成式目标匹配(Generative Target Matching, GTM),该技术扩展了目录有效约束解码,通过基于定向规则生成的位掩码和布隆过滤器匹配器,对语义ID(Semantic ID, SID)前缀进行个性化过滤;与单纯的约束解码相比,SID级GTM将最终广告级目标匹配通过率从23.55%提升至40.42%。针对计算成本和延迟,GRACE以比大语言模型(LLM)更轻量的编码器-解码器Transformer为目标,围绕宽波束、短序列场景重新设计解码器,涵盖注意力核、键值缓存(KV cache)和波束搜索优化。在NVIDIA GH200平台上,与FlashAttention-2和FlashAttention-3基线中更快的一方相比,GRACE在解码步骤中将交叉注意力延迟提升68.0倍,自注意力延迟提升23.4-25.8倍;综合以上优化,解码器延迟降低11.1倍,使广告生成式检索满足延迟和计算要求。

英文摘要

Productionizing generative recommenders for high-volume, real-time ads retrieval creates two serving challenges: eligibility, ensuring that each generated ad is eligible for the request under the advertiser's audience targeting rules, and compute, which requires meeting strict latency and GPU cost requirements while remaining capable of generating thousands of ads per request with wide-beam decoding. This paper presents GRACE, a serving system for ads generative retrieval that addresses both challenges. For eligibility, GRACE introduces Generative Target Matching (GTM), which extends catalog-valid constrained decoding with personalized filtering over Semantic ID (SID) prefixes using bitmask and Bloom filter matchers derived from targeting rules. SID-level GTM improves final ad-level target matching pass rate from 23.55% to 40.42% over constrained decoding alone. For compute-cost and latency, GRACE targets encoder-decoder Transformers, which are more lightweight than LLMs. It redesigns the decoder around the wide-beam, short-sequence regime, covering attention kernels, KV cache, and beam search optimizations. On NVIDIA GH200, compared with the faster of FlashAttention-2 and FlashAttention-3 baselines, GRACE improves cross-attention latency by 68.0 times and self-attention latency by 23.4-25.8 times across decode steps. Together, these changes reduce decoder latency by 11.1 times, keeping ads generative retrieval within latency and compute requirements.

补充信息

↑