arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于矩阵分解MDP的候选生成规划

Planning over Matrix-Factorization MDPs for Candidate Generation

Mikhail Trapeznikov, Maksim Utushkin

arXiv 2607.02115首次发表:更新:

AI 中文总结

将推荐系统的用户旅程建模为MDP,通过折叠更新用户状态进行规划,实验表明单步前瞻即可显著提升固定嵌入下的检索效果。

AI 中文摘要

对于推荐服务,我们将用户旅程视为一系列物品推荐:一个有用的物品会改变用户的状态,从而影响下一步检索的内容。标准的矩阵分解检索忽略了这一点——它构建一个用户向量,并根据静态得分返回前$K$个物品,将它们视为独立。我们提出一个狭窄的问题:何时值得对折叠更新引起的用户状态动态进行规划?为了回答这个问题,我们提出将前$K$检索建模为隐式ALS后验$(A^{-1},u)$上的MDP,其中动作是物品,转移是闭式秩一折叠更新,轨迹奖励结合了相关性相似度和后验对齐项。在相同的固定嵌入下,我们比较了静态检索、单步规划和horizon-$K$ MCTS在五个数据集和两种协议上的表现:每个用户的留最后$n$划分和更严格的全局时间划分。在留最后$n$划分下,动态感知规划在所有数据集上倾向于优于静态检索,并且在全局时间划分下,在MovieLens-1M和VK-LSVD切片上保持了增益。单步前瞻已经捕获了大部分增益,因此轻量级规划层将静态前$K$评分转化为短决策,并改进了固定协同过滤嵌入上的检索,无需重新训练或改变表示。这些增益依赖于使用余弦而非内积相似度来衡量相关性,否则会与物品流行度纠缠。

英文摘要

For a recommender service, we view the customer journey as a chain of item recommendations: a useful item changes the user's state and therefore what should be retrieved next. Standard matrix-factorization retrieval ignores this -- it builds one user vector and returns the top-$K$ items by a static score, treating them as independent. We ask a narrow question: when is it worth planning over the user-state dynamics that fold-in induces? To answer it we propose casting top-$K$ retrieval as an MDP over the implicit-ALS posterior $(A^{-1},u)$, where an action is an item and the transition is a closed-form rank-one fold-in, and the trajectory reward combines a relevance similarity with a posterior-alignment term. Under the same fixed embeddings we compare static retrieval, one-step planning, and horizon-$K$ MCTS across five datasets and two protocols: a per-user leave-last-$n$ split and a stricter global time split. Dynamics-aware planning tends to overcome static retrieval on all datasets under leave-last-$n$, and the gains hold on MovieLens-1M and the VK-LSVD slices under the global time split. A single step of lookahead already captures most of the gain, so the lightweight planning layer turns static top-$K$ scoring into a short decision and improves retrieval over fixed collaborative-filtering embeddings, with no retraining and no change to the representation. These gains depend on measuring relevance with cosine rather than inner-product similarity, which is otherwise entangled with item popularity.

CommentsAccepted to the 5th Workshop on End-to-End Customer Journey Optimization at KDD 2026. 6 pages, 3 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑