arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VaLiDRec:用于生成式推荐的可变长度大语言模型对齐语义ID

VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation

Shutong Qiao, Wei Yuan, Tong Chen, Hao Wang, Quoc Viet Hung Nguyen, Hongzhi Yin

arXiv 2607.25209首次发表:更新:

发表机构

University of Queensland; Computer Network Information Center, Chinese Academy of Sciences; Griffith University(昆士兰大学; 中国科学院计算机网络信息中心; 格里菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对生成式推荐中固定长度语义标识符的问题,提出VaLiDRec框架,通过令牌重要性估计等构建可变长度、大语言模型对齐语义标识符,纳入图感知软提示建模用户偏好,实验证明其性能优越,为生成式推荐提供新范式。

AI 中文摘要

生成式推荐通常使用通过聚类和量化构建的固定长度语义标识符(SID)来表示项目。然而,这些人工编码可能会过度压缩项目语义,与预训练的大语言模型词汇表不一致,并且需要昂贵的自回归解码。鉴于此,我们提出了VaLiDRec,一个基于可变长度、大语言模型对齐语义标识符的生成式推荐框架。VaLiDRec通过令牌重要性估计、语义质量感知修剪和冲突感知细化,直接从信息丰富的原生大语言模型词汇令牌构建SID,使标识符长度能够适应项目语义复杂性。为了对用户偏好进行建模,VaLiDRec纳入了图感知软提示,并将推荐重新表述为具有令牌级项目评分的令牌集预测,消除了自回归SID生成和束搜索。在四个真实世界数据集上的实验表明,VaLiDRec在所有评估指标上始终优于强大的序列和生成式推荐基线。它还实现了卓越的零样本项目冷启动性能,推理速度比LC-Rec快87.49倍。这些结果表明,大语言模型原生的可变长度语义标识符为生成式推荐提供了一种更具表现力和效率的范式。

英文摘要

Generative recommendation commonly represents items using fixed-length semantic identifiers (SIDs) constructed through clustering and quantization. However, these artificial codes may overcompress item semantics, remain misaligned with pretrained LLM vocabularies, and require costly autoregressive decoding. In light of this, we propose VaLiDRec, a generative recommendation framework based on variable-length, LLM-aligned semantic identifiers. VaLiDRec constructs SIDs directly from informative native LLM vocabulary tokens via token importance estimation, semantic-quality-aware pruning, and collision-aware refinement, allowing identifier lengths to adapt to item semantic complexity. To model user preferences, VaLiDRec incorporates graph-aware soft prompts and reformulates recommendation as token-set prediction with token-level item scoring, eliminating autoregressive SID generation and beam search. Experiments on four real-world datasets show that VaLiDRec consistently outperforms strong sequential and generative recommendation baselines across all evaluation metrics. It further achieves superior zero-shot item cold-start performance and 87.49$\times$ faster inference than LC-Rec. These results demonstrate that LLM-native variable-length semantic identifiers provide a more expressive and efficient paradigm for generative recommendation.

Comments10 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑