arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

电商场景中用于点击率预测的原生多模态表示学习

Native Multimodal Representation Learning for Click-Through Rate Prediction in E-Commerce Scenarios

Chao Yi, Feifan Yang, Jiawei Feng, Sishuo Chen, Zhangming Chan, Xiang-Rong Sheng, Han Zhu

arXiv 2608.24091首次发表:更新:

发表机构

Taobao & Tmall Group of Alibaba; University of Science and Technology of China(阿里巴巴淘宝天猫集团; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有多模态CTR预测模型两阶段范式的缺陷,本文提出Mine-Then-Train方法,通过挖掘CTR数据的高质量样本微调多模态编码器,提升了模型性能。

AI 中文摘要

多模态表示已被广泛应用于工业级电商推荐系统中,凭借其强大的语义理解与泛化能力,可提升传统基于稀疏ID的点击率(CTR)预测模型的性能。当前CTR预测任务中的多模态应用框架通常遵循两阶段范式:首先在特定推荐场景的数据上预训练多模态编码器;再利用该预训练多模态编码器提取物品的多模态表示,并将其整合至CTR预测模型中。然而,多模态预训练任务的训练目标与数据分布常与CTR任务存在差异,这限制了多模态表示在下游任务中的有效性。本文聚焦于如何为CTR预测任务学习原生多模态表示,一种直观的解决方案是在CTR任务上联训练多模态编码器与CTR模型,期望编码器能自动学习与下游相关的知识,但我们发现端到端训练并未为现有多模态应用范式带来性能提升。我们的分析显示,原始CTR数据中的用户行为由多模态语义与非多模态因素共同驱动,导致监督信号模糊、编码器更新不一致。为解决该问题,我们提出了Mine-Then-Train方法,该方法从CTR数据中挖掘高质量、具备多模态可解释性的训练样本,并利用这些样本微调多模态编码器,使其更好地适配用户点击偏好。离线与在线实验均验证了该方法的有效性。

英文摘要

Multimodal representations have been widely adopted in industrial e-commerce recommendation systems. Due to their strong semantic understanding and generalization capabilities, they enhance the performance of traditional sparse ID-based Click-Through Rate (CTR) prediction models. Current multimodal application frameworks in the CTR prediction task typically follow a two-stage paradigm: first, pre-training a multimodal encoder on data from specific recommendation scenarios; second, extracting items' multimodal representations using this pre-trained multimodal encoder and integrating them into the CTR prediction model. However, the training objectives and data distribution of multimodal pre-training tasks often differ from those of the CTR prediction task, which limits the effectiveness of multimodal representation on downstream tasks. In this paper, we focus on how to learn Native Multimodal Representation for the CTR prediction task. One intuitive solution is to jointly train the multimodal encoder and CTR model end-to-end on the CTR task, with the expectation that the encoder can automatically learn downstream-relevant knowledge. However, we find that the end-to-end training does not bring performance improvements to existing multimodal application paradigms. Our analysis reveals that user behaviors in raw CTR data are driven by both multimodal semantics and non-multimodal factors, leading to ambiguous supervision and inconsistent encoder updates. To address this, we propose a Mine-Then-Train method that mines high-quality, multimodally interpretable training samples from CTR data and uses them to fine-tune the multimodal encoder for better alignment with user click preferences. Offline and online experiments demonstrate the effectiveness of our approach.

CommentsAccepted at CIKM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑