arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10357cs.LG

后续测试集并非新领域:预训练熟悉度在无污染留出集上依然存在

A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out

  • BrightMind AI
  • University of Texas at Arlington(德克萨斯大学阿灵顿分校)

机构由 AI 辅助整理,请以论文原文为准。

Mahdi Naser Moghadasi, Faezeh Ghaderi

AI总结:

通过构建无污染时间留出集,发现预训练模型优势源于语料库熟悉度而非记忆,提出基准需按语料库声明领域留出集。

AI中文摘要:

时间序列基础模型几乎完全在早于其发布的公共档案上进行评估,因此高分无法与预训练期间见过测试集区分开来。显而易见的补救措施是使用晚于模型发布的留出集。我们构建了这样一个留出集:涵盖来自五个领域的七组数据,包含十三个预测器——四个经典方法、四个按数据集训练的方法、六个预训练方法——所有观测值均在最后一个模型发布之后发表,且每个数据集无需API密钥即可重建。在此协议下,预训练模型赢得了七组中的五组,输给了一个Theta基线一组,并且在每日汇率上,与季节性朴素预测以及所有其他测试方法无法区分。随后我们探究了胜负之间的差异,并报告了一个负面结果:人们会想到的两个内在属性——在输入窗口上测量的季节强度和谱熵——并不能解释这种模式,季节强度甚至与优势呈负相关。真正与之相关的是语料库熟悉度。我们最大的收益(在每周维基百科页面浏览量上,比最佳经典方法低28%的MASE)出现在维基百科页面浏览量上,这是TimesFM作者描述为其预训练语料库主要组成部分的领域,粒度相同,仅时间窗口不同。在预训练模型家族内部,由于每个模型预测相同的序列,序列难度被抵消,TimesFM家族在维基百科上的排名比Chronos家族高出-0.53个排名,而在其他所有地方为-0.09(1500对754个序列,Mann-Whitney p < 1e-5)。我们得出结论:时间留出集消除了对窗口的记忆,但并未消除对领域的熟悉度;因此,基准测试需要相对于已披露的语料库来声明领域留出集;从业者的问题与其说是哪个模型更好,不如说是他们的领域是否是模型训练时接触过的领域。

英文摘要:

Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious remedy is a hold-out that postdates the models. We build one: thirteen forecasters -- four classical, three trained per dataset, six pretrained -- on seven groups drawn from five domains, every observation published after the last model was released, and every dataset rebuildable without an API key. Under this protocol pretrained models win 5 of 7 groups, lose one to a Theta baseline, and on daily exchange rates are indistinguishable from a seasonal naive forecast, along with every other method tested. We then ask what separates the wins from the losses, and report a negative result: the two intrinsic properties one would reach for -- seasonal strength and spectral entropy, measured on the input window -- do not account for the pattern, and seasonal strength is if anything negatively associated with the advantage. What does track it is corpus familiarity. Our largest gain (28% lower MASE than the best classical method, on weekly Wikipedia pageviews) falls on Wikipedia pageviews, the domain TimesFM's authors describe as the bulk of its pretraining corpus, at the same granularities and differing only in time window. Within the pretrained family, where every model forecasts identical series so that series difficulty cancels, the TimesFM family outranks the Chronos family by -0.53 ranks on Wikipedia against -0.09 everywhere else (1,500 vs. 754 series, Mann-Whitney p < 1e-5). We conclude that a temporal hold-out removes memorisation of a window but not familiarity with a domain, that benchmarks therefore need domain hold-outs stated relative to disclosed corpora, and that the practitioner's question is less which model is better than whether their domain is one the model was raised on.

补充信息

↑