arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格数值拉伸变换

Tabular Numeric Stretch Transformation

Zihao Ye, Juyong Kim, Johnna Sundberg, Burak Varici, Pradeep Ravikumar

arXiv 2608.09162首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对表格数值特征变换研究不足的问题,提出含无监督与有监督变体的拉伸变换框架,经38个数据集实验,其有监督变体性能优于所有基线,为表格深度学习提供新策略。

AI 中文摘要

表格数据因异构性给深度学习带来独特挑战,其中数值特征呈现多样的分布、尺度和统计特性。尽管近期进展提升了模型从表格数据学习的能力,但数值数据如何转换为模型友好表示的问题仍未得到充分探索。我们提出拉伸变换框架,该框架将数值特征预处理表述为优化问题,以使得目标函数更平滑从而更易于学习。我们的框架有两个变体:(1)无监督拉伸,通过极小极大优化均匀重新分布特征密度;(2)有监督拉伸,从目标函数平滑性的角度优化目标感知数值特征变换,方法是最小化变换空间中目标函数的狄利克雷能量。我们的理论分析进一步将该框架与几种流行变换关联:无监督拉伸通过共享分段线性几何与分段线性编码密切相关,且当分箱数量增加时趋近于经验CDF变换;而有监督拉伸在细分箱极限下与目标编码密切相关。在TALENT基准的38个数据集上开展的综合实验表明,有监督拉伸始终优于所有基线方法。这些结果表明,针对目标函数平滑性进行显式优化是表格深度学习中一种强大且未被充分探索的策略。

英文摘要

Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scales, and statistical properties. Although recent advances have improved how models learn from tabular data, how numeric data are transformed into model-friendly representations remains comparatively underexplored. We introduce the stretch transformation framework, which formulates numeric feature preprocessing as an optimization problem to make the target function smoother and thus more learnable. Our framework has two variants: (1) unsupervised stretch, which uniformly redistributes feature density via minimax optimization, and (2) supervised stretch, which optimizes target-aware numeric feature transformations from the perspective of target-function smoothness by minimizing the target function's Dirichlet energy in the transformed space. Our theoretical analysis further connects this framework to several popular transformations: unsupervised stretch is closely related to Piecewise Linear Encoding through a shared piecewise-linear geometry and approaches the empirical CDF transformation as the number of bins grows, while supervised stretch becomes closely related to target encoding in the fine-binning limit. Comprehensive experiments on 38 datasets from the TALENT benchmark demonstrate that supervised stretch consistently outperforms all baselines. These results show that explicitly optimizing for target function smoothness is a powerful and underexplored strategy for tabular deep learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑