ScalarLens:具有稳定坐标和上下文响应的数值嵌入用于点击率预测
ScalarLens: Numerical Embeddings with Stable Coordinates and Contextual Responses for CTR Prediction
浏览论文内容
中文总结 AI 辅助
ScalarLens提出一种数值嵌入方法,通过稳定坐标和上下文响应解耦数值位置与含义,在CTR预测中显著优于现有方法。
中文摘要 AI 辅助
用于点击率(CTR)预测的数值嵌入建立在一个方便但受限的前提之上:一个标量只有一个表示。该前提将数值所在的位置与其对当前样本的意义混为一谈。在Criteo验证集划分上,即使去除了加性主效应后,相同的数值区间在类别和数值上下文中仍携带符号相反的残余点击证据。生产流水线加剧了这种不匹配,因为外部归一化的特征需要转换和统计信息在训练和推理之间保持同步。我们引入了ScalarLens,一种数值嵌入,它保留了数值是什么,同时调整了应如何解释它。一个单调局部网格仅从焦点标量构建稳定坐标;有界低秩动态随后产生上下文响应,而不移动该坐标或替换类别标记和CTR主干。在包含19种表示、三个数据集、九个主干和三个种子的1539次运行的主要评估中,ScalarLens在原始数值尺度上的27个设置中排名第一25次,在其余两个设置中排名第二。匹配的消融实验表明,尺度校正、额外局部容量和通用条件化无法复现该增益。一项受控研究进一步在上下文偏移下恢复了类别、数值和混合响应机制,同时焦点坐标保持完全不变。在共享标准化下的完整重跑相比DEER、DAES和NaryDis仍保留显著优势,表明该结果不能仅由对原始尺度的容忍来解释。因此,ScalarLens将数值嵌入重新定义为测量问题:坐标属于数值,而预测响应属于上下文中的数值。
英文摘要
Numerical embeddings for click-through rate (CTR) prediction are built on a convenient but restrictive premise: a scalar has one representation. This premise conflates where a value lies with what it means for the current sample. On the Criteo validation split, the same numerical interval carries residual click evidence with opposite signs across categorical and numerical contexts, even after additive main effects are removed. Production pipelines compound this mismatch because externally normalized features require transformations and statistics to remain synchronized between training and serving. We introduce ScalarLens, a numerical embedding that preserves what a value is while adapting how it should be interpreted. A monotone local mesh constructs a stable coordinate from the focal scalar alone; bounded low-rank dynamics then produce a contextual response without moving that coordinate or replacing categorical tokens and the CTR backbone. In a 1,539-run primary evaluation covering 19 representations, three datasets, nine backbones, and three seeds, ScalarLens ranks first in 25 of 27 settings on original numerical scales and second in the remaining two. Matched ablations show that scale correction, additional local capacity, and generic conditioning do not reproduce the gain. A controlled study further recovers categorical, numerical, and mixed response mechanisms under context shift while the focal coordinate remains exactly invariant. A complete rerun under shared standardization retains significant advantages over DEER, DAES, and NaryDis, showing that the result is not explained by tolerance to raw scales alone. ScalarLens therefore recasts numerical embedding as a measurement problem: coordinates belong to values, while predictive responses belong to values in context.
发表机构
- Ant Group(蚂蚁集团)
- Alibaba Inc.(阿里巴巴集团)
- Henan Polytechnic University(河南理工大学)
机构由 AI 辅助整理,请以论文原文为准。