AI 中文总结
本文针对多目标工业推荐的碎片化问题,提出可控生成式检索框架Multi-Decoder OneRec,通过多解码器结构实现共享建模与目标控制,在公开基准Kwai26及生产环境测试中均取得推荐效果提升。
AI 中文摘要
工业推荐系统通过为特定目标的检索路由分配明确配额来构建候选池,该设计可实现配额控制,但随着路由集规模增长,建模、训练与 serving 过程的碎片化问题日益凸显。基于语义ID的生成式检索提供了一种统一的替代方案,但单一解码器会使目标策略相互纠缠,限制候选的互补性。为此,本文提出Multi-Decoder OneRec,这是一种可控框架,结合了共享表示、独立目标适配与协同解码:所有目标共享用户上下文模块与通用解码器,每个目标额外添加一个独立的、参数高效的LoRA专家。训练阶段,曝光样本的下一个token预测(NTP)更新共享基础,目标过滤后的NTP更新基于事件的专家,Kullback-Leibler(KL)正则化策略优化更新观看时长专家;梯度路由隔离这些更新,通用解码器提供停止梯度参考。推理阶段,明确的路由配额分配固定预算,多解码器约束束搜索减少跨路由重叠。本文公开发布了Kwai26,这是一个大规模多目标基准,包含13.1亿条原始物品级记录、3185万条物品ID条目、2503万个具有有效语义ID的物品,以及预定义划分与评估协议。在相同的512项检索预算下,Multi-Decoder OneRec在四个Recall@512指标上较单一解码器OneRec基线提升了1.69%-5.62%;在生产环境A/B测试中,其单设备应用使用时长相对提升0.37%、第7天留存用户相对提升0.19%、至少有一次分享的设备相对提升0.19%、新内容冷启动相对提升2.09%。这些结果表明,生成式检索可将共享建模与目标特定控制及互补候选生成相结合。
英文摘要
Industrial recommender systems build candidate pools by assigning explicit quotas to objective-specific retrieval routes. This design offers quota control but increasingly fragments modeling, training, and serving as the route set grows. Semantic-ID-based generative retrieval provides a unified alternative, yet a single decoder entangles objective policies and limits candidate complementarity. We propose Multi-Decoder OneRec, a controllable framework that combines shared representations, isolated objective adaptation, and coordinated decoding. All objectives share a user-context module and the General Decoder, while each objective adds an isolated, parameter-efficient LoRA expert. During training, exposure-sample next-token prediction (NTP) updates the shared base, target-filtered NTP updates the event-based experts, and Kullback-Leibler (KL)-regularized policy optimization updates the Watch-time expert; gradient routing isolates these updates, and the General Decoder supplies a stop-gradient reference. At inference, explicit route quotas allocate the fixed budget and Multi-Decoder Constrained Beam Search reduces cross-route overlap. We publicly release Kwai26, a large-scale multi-objective benchmark with 1.31 billion raw item-level records, 31.85 million Item-ID entries, and 25.03 million items with valid Semantic IDs, together with predefined splits and an evaluation protocol. Under the same 512-item retrieval budget, Multi-Decoder OneRec improves over the single-decoder OneRec baseline by 1.69%-5.62% across four Recall@512 metrics. In a production A/B test, it yields relative gains of 0.37% in app usage time per device, 0.19% in Day-7 retained users, 0.19% in devices with at least one share, and 2.09% in new-content Cold-Start. These results show that generative retrieval can combine shared modeling with objective-specific control and complementary candidate generation.
Comments9 pages, 4 figures, 11 tables, 2 algorithms