发表机构
EURECOM(欧洲通信研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建无数据泄露的延迟感知测试框架,审计TAFAS等四种TTA方法在五个基准数据集上的性能,发现多数方法在因果延迟标签下增益有限,部分方法性能甚至差于冻结模型,RLS滤波器组成本低且性能优异。
AI 中文摘要
时间序列预测的测试时自适应(TTA)方法会根据传入的真实标签更新部署的模型或其周围的小型适配器。但H步预测的标签仅在H步后才会出现,而实际数据管道还会增加额外延迟。我们构建了一个无数据泄露的测试框架,其中预测原点s的标签仅在s+d步(d≥H)时才会被释放用于更新,我们在四个近期TTA方法(TAFAS、COSA、PETSA和DynaTTA)的发布代码中强制执行该规则,在五个基准数据集(ETTm1、ETTh2、Weather、Electricity和Traffic)上运行这些方法及其各自的骨干网络和检查点。作为参考,我们添加了两种闭式校正器:一种是由逐坐标中位数组合的递归最小二乘(RLS)滤波器组,无可调超参数,在7通道流上每步耗时56微秒;另一种是ELF式线性校正器。在因果延迟标签下,结果呈现不对称性:在ETTm1上,所有经审计的方法均实现了真正的自适应,但RLS滤波器组的成本仅为其中三种方法的一小部分,性能却优于它们;仅DynaTTA在最小因果延迟下优于该滤波器组,但其每更新一次的成本约为滤波器组的2500倍;ELF式校正器的性能优于所有四种方法。在其余四个数据集上,任何已发表方法相对于其自身冻结检查点实现的最大统计显著改进仅为0.5%;在所有四个数据集上,至少有一种已发表方法在最小因果延迟下的性能显著差于冻结模型;在漂移严重的ETTh2上,更长的标签延迟会使所有与冻结模型分离的适配器(包括我们的适配器)产生显著危害。存在数据泄露的下一步更新会使简单适配器的表观增益最多膨胀110%,而骨干网络训练配方会使冻结在线误差变化多达25倍,超过我们测量到的任何自适应效应。我们发布了该测试框架、集成补丁和所有缓存运行结果。
英文摘要
Test-time adaptation (TTA) methods for time-series forecasting update a deployed model, or a small adapter around it, from incoming ground truth. But the label of an $H$-step forecast exists only $H$ steps later, and real data pipelines add further delay. We build a leakage-free harness in which the label of forecast origin $s$ is released for updates only at step $s+d$ with $d \ge H$, and enforce this rule inside the released code of four recent TTA methods (TAFAS, COSA, PETSA and DynaTTA), run on their own backbones and checkpoints across five benchmarks (ETTm1, ETTh2, Weather, Electricity and Traffic). As references we add two closed-form correctors: a bank of recursive least squares (RLS) filters combined by a per-coordinate median, with no tunable hyperparameters and 56 microseconds per step on the 7-channel streams, and an ELF-style linear corrector. Under causal delayed labels the picture is asymmetric. On ETTm1 every audited method genuinely adapts, yet the RLS bank still beats three of the four at a fraction of their cost; only DynaTTA beats the bank, only at the minimum causal delay, and at roughly 2,500 times the per-update cost; the ELF-style corrector beats all four. On the other four datasets, the largest statistically significant improvement any published method achieves over its own frozen checkpoint is half a percent, on all four at least one published method is significantly worse than the frozen model at the minimum causal delay, and on drift-heavy ETTh2 longer label delays make every adapter that separates from the frozen model, ours included, significantly harmful. Leaky next-step updates inflate the apparent gains of simple adapters by up to 110%, and the backbone training recipe moves frozen online error by up to a factor of 25, more than any adaptation effect we measure. We release the harness, integration patches and all cached runs.
Comments34 pages, 9 figures. Code and cached results: https://github.com/Xodios/TRAINING-ON-THE-FUTURE-A-DELAY-AWARE-AUDIT-OF-TEST-TIME-ADAPTATION-FOR-TIME-SERIES-FORECASTING