arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16633cs.IR

超越固定深度和宽度:优化基于大语言模型的生成式推荐中的文本解码树

Beyond Fixed Depths and Widths: Optimizing Textual Decoding Tries in LLM-based Generative Recommendation

Jingzhe Liu, Hanbing Wang, Jiliang Tang, Liam Collins, Tong Zhao, Neil Shah, Mingxuan Ju

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于大语言模型的生成式推荐中解码树结构问题,提出BONSAI框架,从项目元数据提取信息丰富的词,用最小集覆盖公式构建满足自适应ID长度和约束分支因子属性的解码树,实验显示比基线有高达21.6%的相对改进。

中文摘要 AI 辅助

生成式推荐(GR)在推荐系统中越来越流行,有一系列突出的工作使用大语言模型作为自回归主干来预测下一个项目的词元ID(如标题或关键词)。自回归生成的成功取决于在解码树上进行约束束搜索,以确保生成的输出对应有效项目。然而,当前研究主要集中在生成更全面的词元ID来描述项目,而很大程度上忽略了由这些词元形成的解码树的结构设计,这可能导致树不适用于束搜索,从而降低性能。为解决此问题,我们从解码树优化的角度研究词元ID的有效性。通过实证和理论分析,我们确定了高性能树的两个理想属性:(1)自适应和可变ID长度,使语义丰富度不同的项目能用适当长度的ID表示;(2)约束分支因子,特别是在浅层,这大幅提高了约束束搜索的成功率。受这些属性启发,我们引入了BONSAI:用于自适应标识符的分支优化节点结构,这是一个共同设计文本词元ID及其底层解码树的新颖框架。BONSAI从项目元数据中提取推荐信息丰富的词,并采用最小集覆盖公式递归构建满足上述属性的树。实验表明,BONSAI比现有基线实现了高达21.6%的相对改进。进一步分析证实了我们提出的属性的关键作用,并证明了它们可推广应用于提高其他词元ID方法的性能。

英文摘要

Generative recommendation (GR) is an increasingly popular paradigm in recommender systems, with a prominent line of work using LLMs as autoregressive backbones to predict the next item's term IDs (e.g., titles or keywords). The success of autoregressive generation hinges on constrained beam search over a decoding trie to ensure that generated outputs correspond to valid items. However, current research predominantly focuses on generating more comprehensive term IDs to describe items, while largely neglecting the structural design of the decoding trie formed by these terms. This can lead to a trie that is poorly suited to beam search, which degrades performance. To address this, we examine the effectiveness of term IDs from the perspective of decoding trie optimization. Through empirical and theoretical analyses, we identify two desirable properties for a highly performant trie: (1) adaptive and variable ID length, enabling items with varying semantic richness to be represented by IDs of appropriate lengths, and (2) constrained branching factors, especially at shallow levels, which drastically improves the success rate of constrained beam search. Motivated by these properties, we introduce BONSAI: Branching-Optimized Node Structure for Adaptive Identifiers, a novel framework that co-designs textual term IDs and their underlying decoding trie. BONSAI extracts recommendation-informative words from item metadata and employs a minimum set cover formulation to recursively build a trie that satisfies the above properties. Experiments reveal that BONSAI achieves up to a 21.6% relative improvement over state-of-the-art baselines. Further analyses confirm the crucial role of our proposed properties, and demonstrate their generalizability to be applied to enhance the performance of other term ID methods.

↑