arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27789cs.IR

从理解到行动:面向生成式推荐的反馈驱动策略发现

From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation

Zhi Chen, Minmao Wang, Xingchen Liu, Haoqiang Liang, Huihuang Lin, Likang Wu, Hongke Zhao, Yulong Wang, Shijie Yi, Fei Pan, Peng Jiang

中文总结 AI 辅助

该研究针对生成式推荐的“理解-行动差距”,提出反馈驱动智能体框架,经蒸馏迁移至轻量模型实现无LLM在线推理,在基准测试与在线A/B测试中均取得推荐效果提升。

中文摘要 AI 辅助

基于语义ID的生成式推荐器可实现高效的下一项推荐,但其项级监督主要捕捉行为共现与局部转移。大型语言模型(LLM)可通过对异构交互历史的推理补充这些模型,以理解用户当前需求,但LLM并非专为推荐特定结果反馈训练,语言层面合理的推理未必能带来有效的推荐决策,我们将这种不匹配称为“理解-行动差距”。据此,我们区分了捕捉用户当前需求的意图知识,与在该需求下指定推荐方向和拒绝边界的策略知识。为弥合该差距,我们提出一种反馈驱动的智能体框架,该框架首先诱导面向任务的意图,再根据其相对于仅意图基线的增量效用发现推荐策略;候选策略通过结果衍生的反馈而非语言合理性进行评估与优化。我们进一步将得到的意图与策略知识通过双空间关系蒸馏迁移至轻量型语义ID生成器的两个潜在token中,支持无LLM的在线推理。在公开基准上的实验显示其较基线有持续改进,大规模在线A/B测试实现了4.506%的收益提升与4.621%的ADVV提升。

英文摘要

Semantic-ID-based generative recommenders enable efficient next-item generation, but their item-level supervision mainly captures behavioral co-occurrence and local transitions. Large language models (LLMs) can complement these models by reasoning over heterogeneous interaction histories to understand the user's current demand. However, LLMs are not inherently trained with recommendation-specific outcome feedback, and linguistically plausible reasoning therefore does not necessarily lead to effective recommendation decisions. We term this mismatch the Understanding-Action Gap. Accordingly, we distinguish intent knowledge, which captures the user's current demand, from policy knowledge, which specifies the recommendation direction and rejection boundary under that demand. To bridge this gap, we propose a feedback-driven agent framework that first induces task-oriented intent and then discovers recommendation policies according to their incremental utility over an intent-only baseline. Candidate policies are evaluated and refined using outcome-derived feedback rather than linguistic plausibility. We further transfer the resulting intent and policy knowledge into two latent tokens of a lightweight Semantic-ID generator through dual-space relational distillation, enabling LLM-free online inference. Experiments on public benchmarks show consistent improvements over baselines, while large-scale online A/B tests achieve gains of 4.506% in Revenue and 4.621% in ADVV.

↑