AI 中文总结
该研究针对推荐重排序的效率与协调问题,提出DIRECTOR传输引导并行重排序框架,经实验验证其在大规模工业推荐场景中性能优于现有基线。
AI 中文摘要
重排序是一个组合决策问题,旨在从特定请求的候选集中选择并排序出高效用的列表。主流的生成式重排序器采用自回归(AR)模型,其逐位构建列表以捕捉位置间依赖关系。然而,在实际的贪心或受限宽度解码下,基于前缀的搜索可能过早剪枝全局有前景的排列,且存在固有的顺序延迟,限制了固定服务预算下的有效搜索空间。非自回归(NAR)替代方案通过位置并行预测缓解了效率瓶颈,但简单的位置分解过于独立处理不同位置,导致位置间协调不足,可能出现重复或冲突的物品选择。为在保留并行效率的同时引入全局结构协调,我们提出了DIRECTOR(Dynamic Index-based RECommendation with Transport-Optimized Retrieval),一种传输引导的并行重排序框架。DIRECTOR将候选物品映射到连续潜在空间,并并行生成针对所有目标位置的请求条件动态检索索引。训练时,它使用熵正则化的最优传输(OT)提供冲突感知的监督;推理时,它直接对相似度矩阵执行全局硬匹配,生成无重复的列表而无需迭代传输。为进一步将生成器与仅返回标量效用的不透明列表级评估器对齐,我们引入了前缀锚定信用分配机制,将全局奖励转换为位置特定的训练信号。大量离线和在线实验表明,DIRECTOR在大规模工业推荐场景中始终优于强大的重排序基线,实现了显著提升。
英文摘要
Reranking is a combinatorial decision problem that aims to select and order a high-utility slate from a request-specific candidate set. A major line of generative rerankers adopts autoregressive (AR) models, which construct the slate one position at a time to capture inter-position dependencies. However, under practical greedy or bounded-width decoding, prefix-based search may prematurely prune globally promising permutations and incurs inherently sequential latency, restricting the effective search space under a fixed serving budget. Non-autoregressive (NAR) alternatives alleviate this efficiency bottleneck through position-parallel prediction, but naive position-wise factorization treats different positions too independently, leading to insufficient cross-position coordination and potentially duplicate or conflicting item selections. To retain parallel efficiency while introducing global structural coordination, we propose Dynamic Index-based RECommendation with Transport-Optimized Retrieval (DIRECTOR), a transport-guided parallel reranking framework. DIRECTOR maps candidate items into a continuous latent space and generates request-conditioned dynamic retrieval indices for all target positions in parallel. During training, it uses entropy-regularized OT to provide conflict-aware supervision; at inference, it directly performs global hard matching on similarity matrix, producing duplicate-free slates without iterative transport. To further align the generator with an opaque list-wise evaluator that returns only a scalar utility, we introduce a prefix-anchored credit assignment mechanism that converts the global reward into position-specific training signals. Extensive offline and online experiments demonstrate that DIRECTOR consistently outperforms strong reranking baselines, achieving significant improvement in large-scale industrial recommendation scenarios.
Comments13 pages, 1 figure