arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于语言模型推荐系统的数值和嵌入特征分词

Tokenizing Numerical and Embedding Features for LLM RecSys

Zhe Xu, Ankit Peshin, Chiyu Zhang, Feng Qi, Johnson Lui, Anil Ramakrishna, Justin Johnson, Carl Hu, Kaushik Rangadurai, Luke Simon

arXiv 2607.10016首次发表:更新:

发表机构

Meta(Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对多数基于大语言模型的推荐器无法利用非文本信号的问题,提出软令牌融合框架,将数值和嵌入特征映射到LLM嵌入空间,在基于共享参数LLM的双塔检索模型中实例化该框架,实验证明该方法有效提升了检索性能。

AI 中文摘要

大语言模型(LLMs)因其强大的序列建模和表示学习能力,越来越多地被用作推荐系统的骨干架构。然而,大多数基于LLM的推荐器主要在离散文本标记上运行,而实际推荐管道还依赖上游特征工程或预训练编码器产生的连续数值特征和密集嵌入特征。这种不匹配限制了基于LLM的模型利用细粒度非文本信号的能力。我们提出了一个软令牌融合框架,将数值和嵌入特征映射到LLM嵌入空间,使异构推荐信号能通过标准令牌接口被使用。我们在基于共享参数LLM的双塔检索模型中实例化该框架,并引入基于交互的融合模块,在将嵌入和数值软令牌插入最终LLM输入之前对其进行细化。在三个亚马逊推荐基准上的实验表明,软令牌融合比基于LLM的基线提高了检索性能,且基于交互的融合比异构软令牌的直接连接更有效。

英文摘要

Large language models (LLMs) are increasingly used as backbone architectures for recommender systems because of their strong sequence modeling and representation learning capabilities. However, most LLM-based recommenders operate primarily on discrete textual tokens, whereas practical recommendation pipelines also rely on continuous numerical features and dense embedding features produced by upstream feature engineering or pretrained encoders. This mismatch limits the ability of LLM-based models to exploit fine-grained non-textual signals. We propose a soft-token fusion framework that maps numerical and embedding features into the LLM embedding space, allowing heterogeneous recommendation signals to be consumed through the standard token interface. We instantiate the framework in a shared-parameter LLM-based two-tower retrieval model and introduce an interaction-based fusion module that refines embedding and numerical soft tokens before they are inserted into the final LLM input. Experiments on three Amazon recommendation benchmarks show that soft-token fusion improves retrieval performance over LLM-based baselines, and that interaction-based fusion is more effective than direct concatenation of heterogeneous soft tokens.

Journal refSecond Tokenization Workshop @ COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑