AI 中文总结
该研究将合成关系数据库生成器PluRel作为RDB-PFN的外部预训练数据源,通过三种课程策略对比,发现模式优先引导策略仅用少量数据即可恢复RDB-PFN近94%的性能,证实早期接触真实模式的有效性。
AI 中文摘要
关系基础模型(RFMs)预训练需要大规模合成关系数据库,但现有方法将数据生成与模型训练流程紧密耦合。本研究探讨通用合成关系数据库生成器PluRel能否作为关系上下文学习器RDB-PFN的外部数据源,RDB-PFN原预训练采用60万任务单表热身阶段后接约180万任务的适应阶段。我们构建转换管道,将PluRel生成的数据库(含外部构建的二分类预测任务)映射为RDB-PFN训练格式,评估三种课程策略:模式优先引导(SCHEMA-GUIDED FIRST,真实模式后接全合成)、全合成(FULLY SYNTHETIC,全程多样合成模式)、模式滞后引导(SCHEMA-GUIDED LAST,全合成后接真实模式)。仅用约5500个关系数据库(约3.3万个任务,较原协议少约55倍任务)且无单表热身,最优策略模式优先引导在1024次上下文学习下,19个真实基准任务平均ROC-AUC达0.6346,恢复已发布RDB-PFN性能的87.6%(原性能为0.7245);在64次上下文学习下,差距缩小至93.8%(0.6116对比0.6517)。结果表明,结合合适课程设计时,外部合成生成器可为RFMs提供有用预训练信号,训练早期接触真实模式比后期模式适应更有效。
英文摘要
Relational Foundation Models (RFMs) require large-scale synthetic relational databases for pretraining, but existing approaches tightly couple data generation with the model training pipeline. We study whether PluRel, a general-purpose synthetic relational database generator, can serve as an external data source for RDB-PFN, a relational in-context learner originally pretrained with a 600K-task single-table warm-up followed by an approximately 1.8M-task adaptation stage. We build a conversion pipeline that maps PluRel-generated databases, including externally constructed binary prediction tasks, into the RDB-PFN training format and evaluate three curriculum strategies: SCHEMA-GUIDED FIRST (real-world schema then fully synthetic), FULLY SYNTHETIC (diverse synthetic schemas throughout), and SCHEMA-GUIDED LAST (fully synthetic then real-world schema). Using only approximately 5,500 relational databases (approximately 33K tasks), roughly 55x fewer tasks than the original protocol, and no single-table warm-up, our best curriculum (SCHEMA-GUIDED FIRST) achieves 0.6346 average ROC-AUC across 19 real benchmark tasks at 1024-shot context, recovering 87.6% of the published RDB-PFN performance (0.7245). At 64-shot context, the gap narrows to 93.8% (0.6116 vs. 0.6517). Our results demonstrate that external synthetic generators can provide useful pretraining signals for RFMs when combined with appropriate curriculum design and that exposure to a real-world schema early in training is substantially more effective than late-stage schema adaptation.
CommentsProceedings of the 2nd ICML on Foundation Models for Structured Data