发表机构
AI Institute, University of Waikato; LIP6, CNRS, Sorbonne Université; GECAD, Polytechnic of Porto(怀卡托大学人工智能研究所; 索邦大学CNRS LIP6实验室; 波尔图理工学院GECAD研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统研究表格基础模型在数据流上的表现,发现其预测性能最优,简单保留最近示例即可,但服务成本高,需提升架构效率。
AI 中文摘要
表格基础模型(TFMs)通过上下文学习在表格基准测试中优于已建立的机器学习模型。基于这一成功,人们越来越有兴趣将其应用于数据流场景,其中数据连续到达并随时间演变。在数据流上,TFM通过更新其上下文而非参数来适应,因此其准确性和成本取决于它保留哪些示例以及重建上下文的频率。因此,我们对数据流上的TFM进行了系统研究,涵盖内存管理、计算成本以及流特定挑战,如概念漂移和延迟标签。我们发现,TFM实现了最高的预测性能,并且简单地保留最近的示例与现有的内存管理技术一样有效。它们在漂移后比流式学习器恢复得更快,并在标签延迟下保持最高准确性。然而,这种准确性伴随着高昂的服务成本,因为几乎不变的上下文在每次预测时都会被重新编码。这些结果表明,架构效率是上下文流学习的前进方向。
英文摘要
Tabular foundation models (TFMs) outperform established machine learning models on tabular benchmarks through in-context learning. Building on this success, interest is growing in applying them to data streams, where data arrive continuously and evolve over time. On a stream, a TFM adapts by updating its context rather than its parameters, so its accuracy and cost depend on which examples it keeps and how often it rebuilds its context. We therefore present a systematic study of TFMs on data streams, covering memory management, computational cost, and stream-specific challenges such as concept drift and delayed labels. We find that TFMs achieve the highest predictive performance and that simply retaining the most recent examples is as effective as existing memory management techniques. They also recover faster than streaming learners after drift and keep the highest accuracy under label delay. This accuracy, however, comes at a high serving cost, since a nearly unchanged context is re-encoded at every prediction. These results point to architectural efficiency as the way forward for in-context stream learning.