arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29000cs.IRcs.LG

PaletteID:用于多模态点击率预测的原型组合语义标识符

PaletteID: Prototype-Composed Semantic Identifiers for Multimodal CTR Prediction

Huanyu Liu, Baining Chen, Hui Liu, Zengyang Li, Ziyi Huang

首次发表
浏览论文内容

中文总结 AI 辅助

提出PaletteID模型,以SQ-DPP构建原型调色板聚合语义相关原型,在两个公开数据集上提升CTR预测性能,对长尾物品增益更显著且标识符分配更鲁棒、语义更可解释。

中文摘要 AI 辅助

多模态信息可提升点击率(CTR)预测的准确率,有效缓解物品冷启动与长尾问题。现有研究通常将预训练多模态嵌入离散化为语义标识符(SIDs),使模型学习推荐的任务特定语义表示,但存在两大局限:一是码本分配无法保留语义相关性,会丢弃原始嵌入空间中的细粒度连续信号;二是残差码路径高度依赖前缀码,限制了分层标识符的有效表示可扩展性。为解决这些问题,我们提出基于原型的语义标识符PaletteID(PID),其灵感源于基于调色板的颜色组合,使用一组紧凑的代表性原型物品作为语义锚点,衔接预训练多模态内容空间与推荐模型。具体而言,我们首先通过语义质量感知确定性点过程(SQ-DPP)构建原型调色板,该过程同时考虑局部内容密度与全局语义多样性;随后,对每个目标物品,PID检索一系列语义相关的原型并将其聚合为丰富的PID表示,实现丰富且互补的语义建模。在两个公开数据集上的大量实验表明,PID可一致提升CTR预测性能,对长尾物品的增益更显著,且与现有残差SID方法相比,能生成更鲁棒的标识符分配,提供更具可解释性的 token 语义。

英文摘要

Multimodal information can improve the accuracy of click-through rate (CTR) prediction and effectively alleviate item cold-start and long-tail problems. Recent studies commonly discretize pretrained multimodal embeddings into semantic identifiers (SIDs), allowing the model to learn task-specific semantic representations for recommendation. However, existing methods still provide limited gains due to two major limitations. First, codebook assignment fails to preserve semantic relevance and discards fine-grained continuous signals in the original embedding space. Second, the residual code paths are highly dependent on prefix codes, which limits the effective representational scalability of hierarchical identifiers. To address these issues, we propose PaletteID (PID), a prototype-based semantic identifier. Inspired by palette-based color composition, PID uses a compact set of representative prototype items as semantic anchors to bridge pretrained multimodal content space and recommendation models. Specifically, we first construct a prototype palette with Semantic Quality-Aware Determinantal Point Process (SQ-DPP), which jointly considers local content density and global semantic diversity. Then, for each target item, PID retrieves a sequence of semantically related prototypes and aggregates them into an informative PID representation, enabling rich and complementary semantic modeling. Extensive experiments on two public datasets demonstrate that PID consistently improves CTR prediction and yields larger gains for long-tail items. PID also produces more robust identifier assignments and provides more interpretable token semantics than existing residual SID methods.

发表机构

  • Huazhong University of Science and Technology(华中科技大学)
  • Central China Normal University(华中师范大学)

机构由 AI 辅助整理,请以论文原文为准。

↑