关系基础模型在上下文学习中的支持集目标泄漏:模型依赖性与评估可靠性
Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability
- SAP
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究揭示关系基础模型在上下文学习中的支持集目标泄漏问题,构建20个受控特征并跨13个任务评估,发现泄漏效应依赖模型与任务,且影响评估可靠性,强调支持/查询信息边界的重要性。
AI中文摘要:
关系上下文学习(ICL)利用带标签的支持示例及其关联的关系上下文来预测新查询的标签。当支持上下文中存在源自目标的特征但查询中不可用时,会产生一种失败模式。我们将此情形研究为支持集目标泄漏。我们构建了20个受控的源自目标的特征,这些特征在信号保真度、表示形式、语义透明度、覆盖率以及零跳、一跳和二跳关系放置上有所不同,并在13个RelBench任务和五种关系ICL配置(这些配置在ICL头部、消息传递深度、预训练队列或关系编码器架构上有所不同)中对其进行评估。我们评估了匹配的0跳、1跳和2跳泄漏设置,以及包含所有20个泄漏列的全泄漏条件。在所测试的配置中,目标表(0跳)和全泄漏产生了与干净评估相比最大的总体偏差,而较高跳数的影响通常较弱,这与时间可达性、采样和聚合保真度相关的有效暴露差异一致。泄漏效应强烈依赖于任务和模型,并且即使总体变化很小,也可能反转模型变体之间的相对结论。对于泄漏检测器,我们在共同的基线子集上比较了基于积分梯度(IG)的筛选方法与互信息(MI)和留一列(LOCO)方法。在高影响的0跳和全泄漏条件下,排序质量最强,但基于检测器的移除并不能一致地恢复干净评估。一项四项任务的rel-salt案例研究进一步显示了来自原始关系模式的原生模式泄漏候选者存在相同的评估问题。这些结果将支持/查询信息边界确定为可靠的关系ICL评估的重要组成部分。
英文摘要:
Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.