arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03057cs.LG

合成关系数据何时教会模型使用关系?从预训练数据到模型行为追踪预测结构

When Does Synthetic Relational Data Teach Models to Use Relations? Tracing Predictive Structure from Pretraining Data to Model Behavior

Shivam Dubey, Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Aditya Tanna, Vinay Kumar Sankarapu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过数据归因方法,发现合成预训练数据中跨表信息的预测必要性促使关系模型学习关系机制,RelDiff因外键链接增益最大而表现最佳,干预实验验证了其因果作用。

中文摘要 AI 辅助

关系基础模型越来越多地在合成数据库上进行预训练,然而下游基准测试几乎没有揭示为什么一个合成语料库比另一个产生更好的模型。特别是,强大的性能可能源于现实的逐行统计,而模型从未学会使用关系结构。我们将此作为一个数据归因问题来研究:合成预训练数据的哪个属性诱导关系计算?使用四个关系Transformer检查点,这些检查点采用相同的架构、初始化、目标函数和计算预算,在由四个关系数据生成器产生的语料库上进行训练,我们将数据的可测量属性追踪到学习到的计算和下游行为。我们假设当跨表信息对于掩码单元预训练目标具有预测必要性时,关系机制就会出现。RelDiff表现出来自外键链接父节点的迄今为止最大的预测增益,其对应模型对未见数据库上的外键干预具有独特的敏感性。这种依赖性在随机初始化控制下依然存在,随着损坏链接比例的增加而单调增长,并定位于串行跨表通路。最后,在下游推理期间破坏相同机制会消除RelDiff在关系任务上的优势,而结构不敏感模型几乎保持不变。这些结果将合成训练数据的属性与学习到的机制联系起来,并通过干预与下游行为联系起来。

英文摘要

Relational foundation models are increasingly pretrained on synthetic databases, yet downstream benchmarks reveal little about why one synthetic corpus produces a better model than another. In particular, strong performance may arise from realistic row-level statistics without the model ever learning to use relational structure. We study this as a data-attribution problem: which property of synthetic pretraining data induces relational computation? Using four Relational Transformer checkpoints trained with the same architecture, initialization, objective, and compute budget on corpora produced by four relational data generators, we trace a measurable property of the data to learned computation and downstream behavior. We hypothesize that relational mechanisms emerge when cross-table information is predictively necessary for the masked-cell pretraining objective. RelDiff exhibits by far the largest predictive gain from foreign-key-linked parents, and its corresponding model is uniquely sensitive to foreign-key interventions on unseen databases. This dependence survives a random-initialization control, grows monotonically with the fraction of corrupted links, and localizes to a serial cross-table pathway. Finally, disrupting the same mechanism during downstream inference removes RelDiff's advantage on relational tasks while leaving structure-insensitive models nearly unchanged. These results connect a property of synthetic training data to a learned mechanism and, through intervention, to downstream behavior.

发表机构

  • Lexsi Labs

机构由 AI 辅助整理,请以论文原文为准。

↑