arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SemPOI-RL:对齐大语言模型语义推理以实现可解释的异地兴趣点序列生成

SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation

Yunqi Liu, Yang Zhang, Ruixing Zhang, Liangzhe Han, Yi Qiao, Tongyu Zhu, Leilei Sun

arXiv 2608.30399首次发表:更新:

发表机构

State Key Laboratory of Complex & Critical Software Environment; Beihang University(复杂与关键软件环境国家重点实验室; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SemPOI-RL框架通过微调LLM、引入SPAM模块并结合推荐导向强化学习,实现LLM语义推理与结构化序列生成的对齐,在两个真实数据集上优于传统推荐器和LLM基线,可生成可解释的异地POI序列。

AI 中文摘要

大语言模型(LLM)具备强大的语义推理与开放生成能力,但将这些能力与结构化序列生成对齐仍具挑战性,这在异地兴趣点(POI)序列生成任务中尤为突出——模型需从用户家乡行为推断可迁移的旅行意图,适配跨城市兴趣漂移,并在结构约束下生成连贯的目的地轨迹。现有方法要么依赖可解释性有限的潜在ID迁移,要么直接使用LLM生成序列,未将推断的语义明确落地到位置感知预测中。为解决这一缺口,本文提出SemPOI-RL框架,将LLM语义推理与结构化序列生成对齐,以实现可解释的异地推荐。具体而言,首先微调LLM,以自然语言作为可解释的语义中间表示,从用户家乡轨迹推断面向目的地的旅行风格;随后引入语义POI对齐模块(SPAM),将这些推断的风格落地到风格条件掩码自动编码器中,用于位置感知轨迹生成;最后应用带推荐导向奖励的强化学习,将LLM生成的风格与下游序列质量对齐。在两个真实世界数据集上的实验表明,SemPOI-RL始终优于传统推荐器与直接LLM基线,同时能在旅行的不同阶段提供可解释的风格归因。代码可在this https URL获取。

英文摘要

Large language models (LLMs) exhibit strong semantic reasoning and open-ended generation abilities, but aligning these abilities with structured sequential generation remains challenging. This challenge is particularly evident in out-of-town (OOT) POI sequence generation, where a model must infer transferable travel intent from a user's hometown behaviors, adapt to cross-city interest drift, and generate a coherent destination trajectory under structural constraints. Existing approaches either rely on latent ID-based transfer with limited interpretability or directly use LLMs for sequence generation without explicitly grounding inferred semantics into position-aware predictions. To address this gap, we propose SemPOI-RL, a framework that aligns LLM semantic reasoning with structured sequence generation for interpretable OOT recommendation. Specifically, we first fine-tune an LLM to infer destination-oriented travel styles from users' hometown trajectories, using natural language as an interpretable semantic intermediate. We then introduce a Semantic POI Alignment Module (SPAM) to ground these inferred styles into a style-conditioned masked autoencoder for position-aware trajectory generation. Finally, we apply reinforcement learning with recommendation-oriented rewards to align LLM-generated styles with downstream sequence quality. Experiments on two real-world datasets show that SemPOI-RL consistently outperforms both traditional recommenders and direct LLM baselines, while providing interpretable style attribution across different phases of a trip. The code is available at https://github.com/Wind-Flipped/SemPOI-RL .

Comments19 pages in total, including 9 pages of main text and 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑