AI 中文总结
本文提出适用于多种模型的事后概念漂移缓解方法NOMADD,通过参数外推提升模型在18个数据集基准上的性能,训练耗时远低于同类最先进方法。
AI 中文摘要
当特征分布随时间变化(称为数据漂移)或特征与结果变量的关系随时间变化(称为概念漂移)时,表格模型的性能会下降。由于标记数据可能无法立即获取,或重新训练模型可能不切实际,这些问题的实时缓解颇具挑战性。尽管已有工具可减少漂移,但它们通常针对神经网络架构定制,且仅调整模型的训练方式。本文提出一种替代的事后方法以减少概念漂移,该方法适用于从树模型、神经网络到表格基础模型的多种类型模型。当用户受限于高模型准确率、受限推理时间或模型规模等约束,需为特定用例选择不同模型时,该新工具尤为实用。我们的算法在每个标记训练周期分别拟合基础模型,将其参数与所有周期汇总得到的单个锚定模型进行对比,通过低秩分解压缩这些参数变化,并采用带阻尼的正则化预测方法外推每个潜在因子。在18个数据集组成的Drift-Resilient TabPFN基准测试中,按照该基准自身的协议和指标评估,该外推方法提升了所有应用其的基础模型的性能,且仅需数秒训练即可达到与最先进的Drift-Resilient TabPFN相当的性能。相比之下,Drift-Resilient TabPFN需要在数百万个合成数据集上进行预训练,耗时约1300个GPU小时,且推理速度慢几个数量级(取决于模型)。在讨论部分,我们探讨了将该工具扩展到其他模态的前景与挑战。
英文摘要
Tabular model performance degrades when feature distributions change over time or the relationship between features and outcome variables change over time, known as data drift and concept drift, respectively. These issues are challenging to mitigate in real time because labeled data may not be immediately available, or re-training a model could be impractical. While tools exist to reduce drift, they are typically bespoke to neural network architectures and adapt how models are trained. In this paper, we offer an alternative post-hoc method to reduce concept drift, which is applicable to a variety of models, from trees to neural networks to tabular foundation models. This new tool is especially useful when constraints, such as high model accuracy, bounded inference time, or model size requires users to choose between different models for their specific use-cases. Our algorithm fits the base model separately on each labeled training period, measures how its parameters evolve against a single anchor model pooled over all of those periods, compresses those changes with a low-rank factorization, and extrapolates each latent factor forward with a damped, regularized forecast. On the 18-dataset Drift-Resilient TabPFN benchmark, evaluated under that benchmark's own protocol and metric, the extrapolation improves every base family it is applied to, and achieves performance competitive with the state-of-the-art Drift-Resilient TabPFN with seconds of training. In contrast, Drift-Resilient TabPFN requires pre-training on millions of synthetic datasets over approximately 1,300 GPU-hours, and is orders of magnitude slower in inference (depending on the model). In the discussion, we explore the promise and challenges of extending this tool to other modalities.