AI 中文总结
研究如何利用关系数据库预测目标列缺失值,对比无参数和参数化编码器,分析标签输入时RDB编码器属性,通过实验验证简单无参数编码器在多基准任务中仍有强大性能。
AI 中文摘要
给定一个存储异构表格信息的关系数据库(RDB),如何预测感兴趣的某个目标列中的缺失(或未来)值?由于企业环境中潜在目标的空间巨大,每次有新的预测任务时,最好避免从头学习新模型。基于特定于RDB的编码器的冻结基础模型提供了一个可行的解决方案,但理想的设计仍然是一个开放问题。一方面,最近有人认为,某些无参数子图编码器与单表基础模型相结合,可以在无需特定于RDB的预训练的情况下实现接近最优的性能。同时,其他当代研究主张使用经过预训练的参数化编码器来利用可观察标签学习特定于任务的表示。为了解决这种模糊性,我们专门分析了RDB编码器在标签作为输入时的属性,证明了可训练编码器参数的潜在功效的局限性。作为实证验证,我们表明,相当简单的无参数编码器在许多相关的基准测试任务中仍然能够表现出强大的性能。
英文摘要
Given a relational database (RDB) storing heterogeneous tabular information, how can we predict missing (or future) values in some target column of interest? As the space of potential targets is vast across enterprise settings, it is preferable to avoid learning a new model from scratch each time there is a new prediction task. Frozen foundation models based on RDB-specific encoders provide a viable solution, but ideal design remains an open question. On the one hand, it has recently been argued that certain parameter-free subgraph encoders combined with single-table foundation models can achieve near SOTA performance, with no RDB-specific pre-training required. Meanwhile, other contemporary studies advocate for parameterized encoders pre-trained to exploit observable labels for learning task-specific representations. To address this ambiguity, we analyze RDB encoder properties specifically when labels are present as inputs, proving limitations on the potential efficacy of trainable encoder parameters. As empirical validation, we demonstrate that considerably simpler parameter-free encoders are still capable of strong performance across many relevant benchmarking tasks.
CommentsICML 2026 Workshop on Foundation Models for Structured Data