发表机构
Ekimetrics; ETH Zurich(Ekimetrics; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对表格基础模型中上下文演示影响不明的问题,提出TICDA方法,利用线性替代模型在单次前向传播中高效归因,兼顾准确性与成本,适用于错误检测、上下文策划与主动学习。
AI 中文摘要
表格基础模型(TFMs)通过在上下文中提供的带标签演示进行条件化,无需任何参数更新即可实现强大的预测性能。然而,单个演示如何影响特定预测仍未被充分理解。这一差距在实践中具有重要意义:上下文通常由任何可用的带标签数据组装而成,可能导致包含错误标签、冗余或低质量的示例,从而降低性能。标准数据归因方法无法迁移到TFM设置:基于重采样的方法(如DemoShapley)需要组合数量的前向传播,而基于梯度的估计器(如影响函数)需要计算训练点对模型参数的影响,而上下文学习从不更新这些参数。我们引入了TICDA,一种直接从在TFM潜在嵌入上训练的线性替代模型测量上下文中每个演示影响的方法,仅需单次前向传播且成本可忽略。我们展示了TICDA在四项任务中提供了与竞争对手相比的最佳折衷:检测标签错误、策划上下文以在降低推理成本的同时保持预测准确性、生成可跨TFM迁移的归因分数,以及支持高效主动学习的采集策略。
英文摘要
Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.