arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

内存层:为推荐模型训练模型内缓存

Memory Layer: Train the In-Model Cache for Recommendation Models

Liangyuan Na, Gufan Yin, Yixin Bao, Xianjie Chen, Justin Lin, Ziheng huang, Xinyuan Zhang, Wen Zhang, Hao Lin, Xiaoheng Mao, Shuo Tang, Min Yu, Lei Chen, Chao yang, Ziliang Zhao, Mengjiao Zhou, Zheng Qi, Dmitry Barablin, Chuo-Yun Yang, Kaustubh Vartak, Tingting Zhang, Arun Kumar Singh

arXiv 2607.25110首次发表:更新:

发表机构

Meta(Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究推荐系统训练与服务中项目表示差异问题,提出内存层联合训练模型内缓存,整合更新路径。在Instagram Reels生产中应用,提升预测覆盖率、嵌入新鲜度,缩小训练-服务差距,降低计算成本。

AI 中文摘要

推荐系统的早期排序阶段会预计算项目嵌入并在模型内缓存,以便在严格的延迟约束内进行评分。由于此缓存在训练循环之外仅在服务时存在,训练和服务使用不同的项目表示,这一结构差异限制了质量并增加了操作脆弱性。我们表明,联合设计训练和服务路径可从源头上消除这种表示差异。我们引入了内存层,这是一种与模型联合训练的模型内键值嵌入缓存:项目塔在训练期间写入嵌入,模型在服务时读取它们,通过构造,这是项目表示的单一真相来源。始终在线的嵌入涵盖尚未缓存的项目,因此每个项目都能得到预测,并且该设计将三个单独的训练器到预测器的更新路径整合为一个独立的管道。在Instagram Reels上投入生产后,内存层将预测覆盖率从96%提高到100%,将嵌入新鲜度从O(5分钟)提高到O(20秒),并将训练-服务归一化熵(NE)差距缩小多达86%,为最新鲜的内容带来超过2倍的召回率,并使冷启动参与度提高5-6%。由于嵌入是在训练期间生成的,系统无需单独的批量评估或发布时重新计算,在中性服务计算成本下,将训练和发布计算成本降低30%。

英文摘要

Early ranking stages in recommendation systems precompute item embeddings and cache them in-model for scoring within strict latency constraints. Because this cache exists only at serving time, outside the training loop, training and serving use different item representations, a structural discrepancy that limits quality and adds operational fragility. We show that co-designing the training and serving paths removes this representation discrepancy at its source. We introduce the memory layer, an in-model key-value embedding cache co-trained with the model: the item tower writes embeddings during training and the model reads them at serving, one source of truth for item representations by construction. Always-on embeddings cover items not yet cached, so every item receives a prediction, and the design consolidates three separate trainer-to-predictor update paths into a single self-contained pipeline. Deployed in production on Instagram Reels, the memory layer raises prediction coverage from 96% to 100%, improves embedding freshness from $O(5\text{ min})$ to $O(20\text{ s})$, and narrows the training-serving Normalized Entropy (NE) gap by up to 86%, yielding over $2\times$ recall for the freshest content and a 5-6% cold start engagement lift. Because embeddings are produced during training, the system needs no separate bulk-evaluation or publish-time recomputation, cutting training-and-publish computational cost by 30% at neutral serving computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑