arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25339cs.IR

SPARC:用于生成式推荐的序列感知渐进式属性路由与压缩框架

SPARC: Sequence-aware Progressive Attribute Routing and Compression Framework for Generative Recommendation

Chang Liu, Changfa Wu, Hui Qian, Binbin Cao, Jian Wu, Yuliang Yan, Han Zhu, Bo Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对生成式推荐中现有离散语义ID问题,提出SPARC框架,通过建模序列依赖、路由多表示到插槽及轻量级交互,在不增输入长度下丰富用户历史表示,实验证明其性能优于基线,改进源于上下文条件信息保留。

中文摘要 AI 辅助

生成式推荐将商品标记为离散语义ID(SID),并根据用户历史SID序列自回归生成目标商品。现有SID虽包含多模态和结构化信息,但通常静态分配且与当前交互上下文无关。在工业场景中,行为包含异构属性,完全展开这些特征会增加输入长度,直接压缩则可能过早丢弃上下文相关信息。我们提出了SPARC,即用于生成式推荐的序列感知渐进式属性路由与压缩框架。SPARC首先对每个字段类型的序列依赖性进行建模,以获得上下文感知字段表示。然后将不同字段的原始、上下文和标识表示路由到多个插槽,在固定容量下保留互补信息。最后,轻量级跨商品交互集成中间令牌并将每个历史商品压缩为单个令牌。按照压缩前上下文化的原则,SPARC在不增加生成主干输入长度的情况下丰富了用户历史表示。在工业淘宝和公共亚马逊数据集上的实验表明,SPARC优于强大的传统和生成式基线。与静态压缩变体的进一步比较表明,SPARC的改进来自上下文条件信息保留,而不仅仅是增加压缩模块的表现力。

英文摘要

Generative recommendation tokenizes items as discrete Semantic IDs (SIDs) and autoregressively generates target items from users' historical SID sequences. Although existing SIDs incorporate multimodal and structured information, they are typically statically assigned and independent of the current interaction context. In industrial scenarios, each behavior also contains heterogeneous attributes, such as category, brand, price, behavior type, and timestamp. Fully expanding these features greatly increases the input length, while directly compressing them into a single representation may prematurely discard context-relevant information. We propose \textbf{SPARC}, \uline{\textbf{S}}equence-aware \uline{\textbf{P}}rogressive \uline{\textbf{A}}ttribute \uline{\textbf{R}}outing and \uline{\textbf{C}}ompression Framework for Generative recommendation. SPARC first models the sequential dependencies of each field type to obtain context-aware field representations. It then routes the original, contextual, and identity representations of different fields into multiple slots to preserve complementary information under a fixed capacity. Finally, lightweight cross-item interaction integrates the intermediate tokens and compresses each historical item into a single token. Following the principle of contextualizing before compression, SPARC enriches user-history representations without increasing the input length of the generative backbone. Experiments on industrial Taobao and public Amazon datasets demonstrate that SPARC outperforms strong conventional and generative baselines. Further comparisons with static compression variants show that the improvement of SPARC comes from context-conditioned information retention rather than merely increasing the expressiveness of the compression module.

↑