arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于深度神经网络解决表内预测问题并使用合成数据进行性能评估

Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data

Xiao Zhao, Daniela Oelke

arXiv 2609.01262首次发表:更新:

发表机构

Offenburg University; Institute for Machine Learning and Analytics(奥芬堡大学; 机器学习与分析研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出表内预测(ITP)问题,采用自监督学习方法,通过新型神经层处理连续特征缺失值,基于合成数据评估MLP、Resnet、Transformer性能,发现注意力结构在合适条件下表现更优。

AI 中文摘要

表格深度学习(TDL)利用神经网络(NN)从表格数据中提取模式。传统TDL方法遵循监督学习范式,其中明确给定目标特征。然而,本研究探索了一种不同的方法,即使用深度神经网络学习给定表格中各列之间的关系。我们研究神经网络是否可以基于给定表格中其余已知列来预测任意选定列的值,将该问题称为表内预测(ITP),其与表格插补方法及TDL的预训练任务略有不同。我们确定了三个潜在使用场景,据我们所知,这些场景在文献中尚未得到广泛研究。我们采用自监督学习方法解决该问题,通过随机选择要掩码的列并将其用作学习目标。本研究聚焦于仅包含连续特征的表格数据集,为处理连续特征中的缺失值,我们提出了一种新型神经层,用于嵌入数值和空值。我们基于预定义的列关系生成合成数据,并使用两种不同机制插入空值,此外还采用了一种适配的掩码策略来创建测试数据。我们使用生成的合成数据评估了三种神经网络架构,即多层感知机(MLP)、残差网络(Resnet)和Transformer的性能,得出结论:当有足够多的训练样本且选择相对较大的嵌入长度时,基于注意力的结构优于另外两种网络。我们强调,这些发现是在受控的合成条件下,针对少量列获得的,因此应将其视为一项初始的、范围较窄的研究,而非对真实世界表格数据上ITP的一般性描述。

英文摘要

Tabular deep learning (TDL) leverages neural networks (NN) to extract patterns from tabular data. Traditional TDL methods follow a supervised learning paradigm, where a target feature is explicitly given. In this work, however, we explore a different approach by employing deep NNs to learn relationships among individual columns within a given table. We investigate whether NNs can predict the values of arbitrarily selected columns in a given table based on the remaining known columns. We call this problem In-Table Prediction (ITB), which is slightly different from table imputation methods and the pretraining task of TDL. Three potential usage scenarios are identified, which, to our best knowledge, have not been extensively studied in the literature. A self-supervised learning approach is applied to address this problem by randomly selecting columns to be masked out and used as learning targets. This work focuses on tabular datasets containing only continuous features. To handle missing values in continuous features, a novel neural layer is proposed to embed both numerical and empty values. Synthetic data is generated based on predefined column relationships, with empty values inserted using two distinct mechanisms. Additionally, an adapted masking strategy is employed to create test data. Performances of three NN architectures, namely MLP, Resnet and Transformer, are evaluated using the generated synthetic data. We conclude that, the attention-based structure outperforms the other two networks, when a sufficiently large number of training examples is available and a relatively large embedding length is chosen. We stress that these findings are obtained under controlled, synthetic conditions with a small number of columns and it should therefore be regarded as an initial, narrowly-scoped investigation rather than a general characterization of ITP on real-world tabular data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑