发表机构
National Taiwan University; Johns Hopkins University(国立台湾大学; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出INTEGER,通过路由令牌、历史重新锚定和行为回放,将生成式推荐扩展到多轮对话,在保持准确性的同时提升交互质量,并在Amazon数据集上显著提升Hit@10。
AI 中文摘要
生成式推荐器从用户的交互历史中解码项目,但无法让用户纠正错过其当前意图的推荐。添加对话是自然的,因为项目和单词共享相同的输出空间,然而训练模型进行对话可能会覆盖其依赖的历史到项目的映射。我们引入了INTEGER(**INTE**ractive **GE**nerative **R**ecommendation,交互式生成推荐),它将生成式推荐扩展到多轮交互,通过一个学习到的路由令牌让模型决定何时推荐,通过历史重新锚定使每个项目同时基于过去行为和对话,以及通过指令数据排练的行为回放来防止适应过程中的遗忘。因此,用户可以在对话中对推荐提供反馈,而推荐仍基于行为历史,并且准确性不会以流畅性为代价。在Amazon Beauty和Toys上,INTEGER在准确性上达到或超过最强基线,同时具有竞争力的对话质量,在Amazon Beauty上将Hit@10提高了13.3%,并且显著优于其起始的生成式推荐器。我们的分析表明,INTEGER学习了朴素适应无法获得的行为,在用户意图明确时进行推荐,并在推荐时刻保持对行为历史的关注。INTEGER还学习了项目空间上的意图无关替换,该替换抑制了被拒绝的项目,但将属性感知反馈作为下一步方向。
英文摘要
Generative recommenders decode items from a user's interaction history, but offer no way for users to correct a recommendation that misses their current intent. Adding conversation is natural since items and words share same output space, yet training the model to converse may overwrite the history-to-item mapping it relies on. We introduce INTEGER (**INTE**ractive **GE**nerative **R**ecommendation), which extends generative recommendation to multi-turn interaction with a learned routing token that lets the model decide when to recommend, history re-anchoring that conditions each item on both past behavior and the dialogue, and behavioral replay with instruction-data rehearsal that prevents forgetting during adaptation. Users can thus give feedback on recommendations within the dialogue, while recommendations stay grounded in behavioral history and accuracy is not traded for fluency. On Amazon Beauty and Toys, INTEGER matches or exceeds the strongest baselines in accuracy with competitive conversation quality, improving Hit@10 by 13.3% on Amazon Beauty, and significantly outperforms the generative recommender it starts from. Our analyses show that INTEGER learns behaviors that naive adaptation fails to acquire, recommending once the user's intent is clear and staying attentive to behavioral history at the moment of recommendation. INTEGER also learns an intent-agnostic replacement over the item space, which suppresses rejected items but points to attribute-aware feedback as the next step.