AI 中文总结
Think2Go框架统一SFT与RL推理,通过优势加权机制校准策略优化,解决现有POI推荐模型的意图捕捉不足问题,提升推荐性能与稳健性。
AI 中文摘要
下一个兴趣点(POI)推荐任务旨在从历史签到数据中挖掘用户行为偏好模式,以提供个性化的下一个目的地建议。现有方法主要依赖浅层上下文信息和人工构造的特征交互来预测下一个POI。然而,用户移动模式固有的稀疏性和复杂性限制了非推理模型捕捉深层意图的计算能力,而大语言模型(LLMs)的表现欠佳,原因在于当语义ID(SIDs)被单独训练时,它们缺乏对SIDs的深入理解。为解决这些局限,我们提出Think2Go,一种新型的生成式下一个POI推荐框架,该框架通过测试时计算缩放增强模型对SID表示的理解,并探索多样的时空模式。我们将监督微调(SFT)和基于强化学习(RL)的推理统一在单一架构中,实现记忆与自适应推理的联合优化,以更好地保留用户行为模式,同时探索多样的用户偏好。为进一步校准自适应推理中的策略优化,我们提出两种优势加权机制:其一为提示认知不确定性,通过核密度方法估计,用于评估查询与用户历史之间的时空周期模式对齐情况,在高认知不确定性下促进更多探索;其二为奖励感知优势缩放,通过将奖励相对于其最大值归一化以适应更新幅度,从而提高训练稳定性并减轻对噪声信号的过拟合。这种联合校准形成了一种隐式课程学习策略,提供细粒度、实例感知的策略更新,防止熵崩溃并支持稳健的探索。
英文摘要
Next Point-of-Interest (POI) recommendation task focuses on mining user behavioral preference patterns from historical check-ins to provide personalized suggestions for the next destination. Existing methods primarily rely on shallow contextual information and handcrafted feature interactions to predict the next POI. However, the inherent sparsity and complexity of user mobility patterns limit the computational capacity of non-reasoning models to capture deep intent, while large language models (LLMs) perform suboptimally because they lack a deep understanding of semantic IDs (SIDs) when SIDs are trained separately. To address these limitations, we propose Think2Go, a novel generative next POI recommendation framework, which enhances the model's comprehension of SID representations and explores diverse spatial-temporal patterns via test-time computational scaling. We unify supervised fine-tuning (SFT) and reinforcement learning (RL)-based reasoning within a single architecture, enabling joint optimization of memorization and adaptive reasoning to better retain user behavior patterns while exploring diverse user preferences. To further calibrate policy optimization in adaptive reasoning, we propose two advantage weighting mechanisms that integrate (1) prompt epistemic uncertainty, estimated via kernel density methods to assess the spatial-temporal periodic pattern alignment between queries and user history, promoting increased exploration under high epistemic uncertainty; and (2) reward-informed advantage scaling, captured by normalizing rewards against their maxima to adapt update magnitudes, thereby improving training stability and mitigating overfitting to noisy signals. This joint calibration forms an implicit curriculum learning strategy, delivering fine-grained, instance-aware policy updates that prevent entropy collapse and support robust exploration.
CommentsAccepted by KDD 2026 Research Track Cycle 1 (Oral presentation)