STAR:面向点击后转化率(PCVR)预测的结构化分词与目标感知兴趣表示
STAR: Structured Tokenization and Target-Aware Interest Representation for PCVR Prediction
AI总结:
针对KDD Cup 2026腾讯UniRec挑战赛,提出STAR框架,结合结构化分词与目标感知兴趣表示,经实验验证其各组件对PCVR预测的AUC有显著提升。
AI中文摘要:
点击后转化率(PCVR)预测是工业推荐系统中的核心排序任务。现代排序模型需同时捕捉异构非序列特征、多行为用户序列以及目标物品感知的用户兴趣,同时需对高基数稀疏特征、缺失值及训练-推理不一致性保持鲁棒性。本文提出STAR(结构化分词与目标感知兴趣表示),这是面向2026年KDD杯腾讯UniRec挑战赛的实用框架。STAR在HyFormer风格多序列骨干网络的基础上,结合结构化特征分词与目标感知兴趣表示,引入高基数信号恢复、显式用户-物品交互令牌、目标感知序列解码,以及受InfoNCE启发的加权用户-物品对比辅助目标。我们通过从保存的训练配置中重构特征重映射表和结构超参数,进一步对齐训练与推理流程。在挑战赛数据集上的实验确定了最可靠提升排序AUC的组件,同时报告LogLoss作为校准诊断指标。主要消融研究显示,时间上下文带来大幅提升,对比对齐、目标感知兴趣编码及高基数序列特征恢复则带来较小但有用的贡献。
英文摘要:
Post-click conversion rate (PCVR) prediction is a core ranking task in industrial recommender systems. Modern ranking models must jointly capture heterogeneous non-sequential features, multi-behavior user sequences, and target-item-aware user interests, while remaining robust to high-cardinality sparse features, missing values, and train-inference inconsistencies. In this paper, we present STAR (Structured Tokenization and Target-Aware Interest Representation), a practical framework for the KDD Cup 2026 Tencent UniRec Challenge. STAR combines structured feature tokenization with target-aware interest representation on top of a HyFormer-style multi-sequence backbone. It introduces high-cardinality signal recovery, explicit user-item interaction tokens, target-aware sequence decoding, and a weighted user-item contrastive auxiliary objective inspired by InfoNCE. We further align the training and inference pipelines by reconstructing feature remapping tables and structural hyperparameters from the saved training configuration. Experiments on the challenge dataset identify the components that most reliably improve ranking AUC, while LogLoss is reported as a calibration diagnostic. The main ablation study shows a large gain from temporal context, with smaller but useful contributions from contrastive alignment, target-aware interest encoding, and high-cardinality sequence feature recovery.