关系基础模型在上下文学习中的支持集目标泄漏:影响、检测与缓解
Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Impact, Detection, and Mitigation
浏览论文内容
中文总结 AI 辅助
本研究提出关系型上下文学习中的支持集目标泄漏问题,构建14种泄漏类型,用积分梯度检测并移除泄漏列,发现目标表泄漏者影响最大且IG可部分恢复性能。
中文摘要 AI 辅助
关系型上下文学习(ICL)根据标记的支持样本及其关联表来条件化预测,当支持集包含查询不可用的目标派生特征时,会产生一种失败模式。我们将此问题定义为支持集目标泄漏,区别于数据集构建、时间划分或表示学习期间的泄漏。在此情境下,目标派生(泄漏者)列仅存在于关系型上下文推理期间的标记支持集中,而查询保持干净。我们构建了14种合成泄漏者类型,对应20列,涵盖具有不同噪声水平、覆盖率、模态、语义透明度和关系距离的代理。我们在保留的RelBench数据库上评估了一个带有ICL头的冻结关系编码器,并使用积分梯度(IG)对可疑列进行排序和移除。我们的结果表明,支持集泄漏的影响因任务和关系距离而异。目标表泄漏者导致最明显的性能下降,而一跳和二跳泄漏者未被模型一致使用。IG在各数据集中对目标表泄漏者进行高排序,并在泄漏影响最大的设置中部分恢复了性能。
英文摘要
Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct from leakage during dataset construction, temporal splitting, or representation learning. Here, the target-derived (leaker) columns are present only in the labeled support set during relational in-context inference, while queries remain clean. We construct 14 synthetic leaker types, corresponding to 20 columns, spanning proxies with different noise levels, coverage, modalities, semantic transparency, and relational distances. We evaluate a frozen relational encoder with an ICL head on held-out RelBench databases and use Integrated Gradients (IG) to rank and remove suspicious columns. Our results show that the effect of support-set leakage varies across tasks and relational distances. Target-table leakers cause the clearest degradation, while one- and two-hop leakers are not consistently used by the model. IG ranks target-table leakers highly across datasets and partially recovers performance in settings where leakage has the largest effect.
发表机构
- SAP(思爱普)
机构由 AI 辅助整理,请以论文原文为准。