发表机构
Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PACE通过在特征评分前插入冻结的表格基础模型列编码器,将特征扩展为上下文嵌入,以轻量代价提升高维表格学习的非线性特征筛选及下游预测性能。
AI 中文摘要
在高维表格学习中,特征筛选提供了一种轻量级、模型无关的方法,在模型拟合前移除不相关特征。然而,直接对原始值进行评分可能无法捕捉非线性或分布结构。我们提出PACE(即插即用上下文嵌入),它在现有特征评分规则之前插入一个冻结的表格基础模型(TFM)列编码器,将每个特征扩展为更高维度的上下文表示。在受控研究中,PACE以仅适度的额外编码成本,改善了复杂非线性依赖的原始空间筛选。这些收益转化为TALENT数据集上的下游预测:PACE-DC将二分类AUC提升0.077,多分类宏AUC提升0.064,在十个学习器上中位归一化RMSE改善0.063。匹配的随机权重和随机特征对照表明,PACE的收益来自预训练结构,而非单纯的维度扩展。PACE进一步在与任务拟合选择器和基于归因方法的对比中实现了有利的性能-时间权衡,将预训练列几何定位为表格学习的可复用上游原语。
英文摘要
In high-dimensional tabular learning, feature screening provides a lightweight, model-agnostic way to remove irrelevant features before model fitting. However, scoring raw values directly can miss nonlinear or distributional structure. We introduce PACE (Plug-and-Play Contextual Embedding), which inserts a frozen tabular foundation model (TFM) column encoder before an existing feature-scoring rule, expanding each feature into a higher-dimensional contextual representation. Across controlled studies, PACE improves raw-space screening of complex nonlinear dependence with only modest additional encoding cost. These gains translate to downstream prediction on TALENT datasets: PACE-DC improves binary AUC by 0.077 and multiclass macro-AUC by 0.064, with a median normalized RMSE improvement of 0.063 across ten learners. Matched random-weight and random-feature controls show that PACE gains from pretrained structure beyond generic dimensional expansion. PACE further achieves favorable performance--time trade-offs against task-fitted selectors and attribution-based methods, positioning pretrained column geometry as a reusable upstream primitive for tabular learning.