DEGR:双探索驱动的生成重排序方法,用于自适应跨请求上下文桥接
DEGR: Dual Exploration-Driven Generative Re-Ranking for Adaptive Cross-Request Context Bridging
浏览论文内容
中文总结 AI 辅助
针对工业推荐系统重排序受上游供给限制的问题,提出DEGR方法,通过双探索驱动的混合优化范式实现自适应跨请求上下文桥接,在京东电商推荐系统中较SOTA方法取得UCTR最高1.22%、PV0.20%的提升。
中文摘要 AI 辅助
在工业推荐系统中,重排序阶段需在序列级优化中平衡业务目标与多样性,同时建模上下文信息。但受限于固定的上游供给,现有方法难以进一步提升效果,尤其在供给质量较低时。为解决该问题,重排序可主动平衡即时价值与探索价值,例如在低质量供给下优先进行探索性曝光,以保留浏览潜力并促进偶然转化。为此,本文提出双探索驱动的生成重排序(DEGR)方法。DEGR采用混合监督-强化探索与优化范式,由探索性奖励模型引导,自适应平衡即时价值与探索价值。该混合优化范式包含三个关键组件:监督学习、探索多样性约束,以及用于偏好优化的自适应奖励加权ORPO。通过这种双探索,生成器最终充当自适应跨请求上下文桥接。离线与在线实验表明,DEGR在京东电商推荐系统中优于SOTA方法,UCTR提升最高达1.22%,PV提升0.20%。
英文摘要
In industrial recommendation systems, the re-ranking stage balances business objectives and diversity for sequence-level optimization while modeling contextual information. However, constrained by fixed upstream supply, existing methods fail to deliver further effectiveness gains, especially under low-quality supply. To overcome this, re-ranking can actively balance immediate and exploratory value, for instance, by prioritizing exploratory exposure under low-quality supply to preserve browsing potential and facilitate serendipitous conversions. Therefore, we propose a Dual Exploration-Driven Generative Re-Ranking (DEGR) method. DEGR adopts a hybrid supervised-reinforcement exploration and optimization paradigm, guided by an exploratory reward model that adaptively balances immediate and exploratory value. The hybrid optimization paradigm integrates three key components: supervised learning, exploration diversity constraint, and adaptive reward-weighted ORPO for preference optimization. Through this dual exploration, the generator ultimately acts as an adaptive cross-request contextual bridge. Offline and online experiments indicate that DEGR outperforms SOTA methods, achieving improvements of up to 1.22% UCTR and 0.20% PV in the JD E-commerce recommendation system.