发表机构
Lexsi Labs(Lexsi Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用合成预训练数据构建可验证的归因测试平台,通过行为条件归因与反事实干预,验证表格基础模型归因的忠实性,并揭示溯源关联与干预效果的不一致性。
AI 中文摘要
训练数据归因旨在识别哪些训练样本塑造了模型行为,然而验证此类主张是困难的,因为因果训练影响很少是可观察的。我们认为,受控的合成预训练使归因在实验上可检验。使用O'PRIOR,一个用于表格基础模型的富含溯源信息的合成任务生成器,我们构建了一个测试平台,其中每个预训练任务都带有关于结构机制、缺失性、混杂、捷径和分布偏移的明确谱系。我们将行为条件归因与反事实重训练和溯源感知干预相结合,以测试任务级忠实性和机制级一致性。在保留的真实任务上,移除最高归因的5%合成任务使平均ROC-AUC下降0.013,而随机移除的下降为0.002±0.004,同时移除最低归因任务使性能提升0.003。在捷径溯源任务中,定向移除产生0.043的效果,而匹配的随机移除为0.016。溯源判别按排名AUROC(0.55-0.62)较为适中,尽管top-k富集显著,揭示溯源关联和干预忠实性不必重合。我们的结果确立了合成溯源作为可验证贡献性归因的受控设置。
英文摘要
Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002$\pm$0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution