课程很重要:基于合成数据的数据高效关系型PFN预训练
Curriculum Matters: Data-Efficient Relational PFN Pretraining with Synthetic Data
浏览论文内容
中文总结 AI 辅助
该研究针对关系型PFN预训练,发现采用PluRel作为合成数据源,渐进式课程设计可大幅减少所需合成数据量,且其性能接近专用关系型管线,凸显课程设计与合成数据多样性的关键作用。
中文摘要 AI 辅助
关系型先验数据拟合网络(PFN),如RDB-PFN,通过在数百万个合成任务上进行预训练,来近似多表关系型数据库上的贝叶斯推理。我们针对该范式研究了三个相互关联的问题:第一,结构不同的合成生成器PluRel能否替代RDB-PFN的先验?第二,向PFN呈现合成数据的顺序对下游性能有多大影响?第三,在引入任何关系数据之前,仅通过单表合成预训练,PFN能获得多少关系推理能力?在所有实验中均使用PluRel作为唯一合成数据源,我们发现:(i)逐步式单表课程(将模式复杂度从7列逐步拓宽至17列)在包含23个任务的表格基准上,仅使用约13300个合成表(比RDB-PFN报告的预热方案少约45倍的单表数据集),即可达到0.703的平均ROC-AUC;而用相同数据一次性训练则降至0.541的ROC-AUC;(ii)仅用约5500个PluRel数据库从头开始训练的关系型课程,在包含19个任务的RelBench/4DBInfer基准上达到0.638的平均ROC-AUC,以少约220倍的关系型合成数据恢复了RDB-PFN报告性能的88%;(iii)未进行任何关系适配,直接在关系基准上评估的单表课程模型,达到0.631,几乎与专用关系型管线的性能相当。综上,这些发现表明,对于关系型PFN预训练而言,课程设计和合成数据多样性可能比特定的关系生成器或原始合成数据规模本身更为重要。
英文摘要
Relational Prior-Data Fitted Networks (PFNs) such as RDB-PFN approximate Bayesian inference over multi-table relational databases by pretraining on millions of synthetic tasks. We investigate three intertwined questions about this paradigm. First, can a structurally different synthetic generator PluRel substitute for RDB-PFN's prior? Second, how much does the order in which synthetic data is presented to the PFN affect downstream performance? Third, how much relational reasoning can a PFN acquire from single-table synthetic pretraining alone, before any relational data is introduced? Using PluRel as the sole synthetic data source across all experiments, we find: (i) a progressive single-table curriculum that gradually widens schema complexity from 7 to 17 columns reaches 0.703 average ROC-AUC on the 23-task tabular benchmark using only approximately 13,300 synthetic tables (approximately 45x fewer single-table datasets than RDB-PFN's reported warm-up recipe), while the same data trained all-at-once collapses to 0.541 ROC-AUC; (ii) a relational curriculum trained from scratch on only approximately 5,500 PluRel databases reaches 0.638 average ROC-AUC on the 19-task RelBench/4DBInfer benchmark, recovering 88% of RDB-PFN's reported performance with approximately 220x less relational synthetic data; and (iii) the single-table curriculum model, evaluated directly on the relational benchmark without any relational adaptation, achieves 0.631, nearly matching the dedicated relational pipeline. Together, these findings suggest that curriculum design and synthetic data diversity may matter more for relational PFN pretraining than the specific relational generator or raw synthetic scale alone.
发表机构
- SAP Labs, LLC.(思爱普实验室有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。