arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推荐世界模型用于未来状态控制

Recommendation World Models for Future-State Control

Jinfeng Xu, Zheyu Chen, Ziyue Peng, Jianheng Tang, Zheng Lin, Jing Yang, Puzhen Wu, Zheng Xing, Victor C. M. Leung

arXiv 2609.30711首次发表:更新:

发表机构

The University of British Columbia; The Hong Kong Polytechnic University; The Hong Kong University of Science and Technology; Peking University; University of Luxemburg; University Malaya; The University of Hong Kong; Shenzhen University(不列颠哥伦比亚大学; 香港理工大学; 香港科技大学; 北京大学; 卢森堡大学; 马来亚大学; 香港大学; 深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出UA-TWM,一种以效用为锚定的世界模型接口,通过构造和评估附近列表动作,在效用约束下选择替代方案,以改善顺序推荐中的未来状态控制,实验证明其能提升Recall@20、NDCG@20等指标。

AI 中文摘要

顺序推荐优化了要排序的项目,而每个展示的列表也会塑造后续的反馈和用户状态。我们研究训练好的排序器如何支持关于这些未来后果的决策。我们引入了UA-TWM,一个以效用为锚定的世界模型接口,它构造附近的列表动作,估计它们与目标相关的后果,并在效用约束下选择一个替代方案。当没有替代方案符合条件时,参考列表作为回退。一个日志回放实例结合了效用和目标增益估计与校准的失败风险预测;一个闭环实例使用一步状态-动作预测,并在观察到反馈后更新其决策。我们在MovieLens-25M和KuaiRand-Pure上跨十二个顺序推荐骨干评估迁移,并在KuaiSim中评估重复的目标导向交互。附加该接口改善了每个匹配的日志骨干的Recall@20、NDCG@20和未来状态对齐。选择消融揭示了激进目标追求中的效用和风险成本,而闭环诊断隔离了动作条件预测的贡献。因此,局部后果建模使得在训练好的顺序排序器周围能够进行目标感知的选择。

英文摘要

Sequential recommendation optimizes which items to rank, while each displayed slate also shapes subsequent feedback and user state. We study how a trained ranker can support decisions about these future consequences. We introduce UA-TWM, a utility-anchored world-model interface that constructs nearby slate actions, estimates their target-relevant consequences, and selects an alternative subject to utility constraints. The reference slate serves as a fallback when no alternative qualifies. A logged-replay instantiation combines utility and target-gain estimates with calibrated failure-risk prediction; a closed-loop instantiation uses one-step state-action prediction and updates its decisions after observed feedback. We evaluate transfer across twelve sequential backbones on MovieLens-25M and KuaiRand-Pure, and repeated target-directed interaction in KuaiSim. Attaching the interface improves Recall@20, NDCG@20, and future-state alignment for every matched logged backbone. Selection ablations reveal the utility and risk costs of aggressive target pursuit, while closed-loop diagnostics isolate the contribution of action-conditioned prediction. Local consequence modeling thus enables target-aware selection around a trained sequential ranker.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑