arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时间序列验证的不可能三角:训练充分性、测试覆盖性与时间因果性之间的守恒律

The Impossible Trinity of Time-Series Validation: A Conservation Law among Training Sufficiency, Test Coverage, and Temporal Causality

Jiayu Li

arXiv 2609.29530首次发表:更新:

AI 中文总结

本研究证明时间序列验证中训练充分性、测试覆盖性与时间因果性无法同时满足,并给出量化守恒律,指出扩展walk-forward是因果验证的帕累托最优方案。

AI 中文摘要

在时间序列上验证模型需要同时满足三个要求:每次训练应使用样本的大部分(充分性),测试集应共同覆盖样本的大部分(覆盖性),且训练数据应位于测试数据之前(因果性)。我们证明这三者无法同时满足,并为每一项给出代价。设α为各折中最小的训练比例,β为测试所覆盖样本的比例,Λ为测试点未来中被用作训练数据的样本比例,δ为测试点到其未来最近训练点的距离。长度为T的样本上的任何方案都满足α+β≤1+Λ和α+min{β,δ/T}≤1,且在β-混合条件下,测试点的泄漏偏差至多为2Mβ_mix(δ)。换言之:超越因果边界α+β=1需要利用未来数据进行训练;这些未来数据必须位于测试点之后(1-α)T的范围内;其危害取决于距离而非数量。因此,扩展的walk-forward正是因果验证的帕累托前沿,k折交叉验证获得最多的未来数据,而带有禁运期的净化k折则以距离为代价,当过程快速遗忘时这种代价低廉,但无法修复非平稳性所要求的因果性部分。在纯噪声上,打乱的5折报告信息系数为+0.32,而连续的5折使用相同数量的未来数据,报告为+0.004。

英文摘要

Validating a model on a time series asks for three things at once: each training run should use most of the sample (sufficiency), the test sets should together cover most of the sample (coverage), and training data should come before test data (causality). We prove that the three cannot be had together and price each one. Let $α$ be the smallest training fraction over folds, $β$ the fraction of the sample covered by tests, $Λ$ the fraction of the sample used as training data from the future of a test point, and $δ$ the distance from a test point to the nearest training point in its future. Every scheme on a sample of length $T$ satisfies $α+β\le 1+Λ$ and $α+\min\{β,δ/T\} \le 1$, and under $β$-mixing the leakage bias at a test point is at most $2Mβ_{\mathrm{mix}}(δ)$. In words: going beyond the causal frontier $α+β=1$ requires training on the future; that future data must sit within $(1-α)T$ of a test point; and its harm depends on its distance, not its amount. Hence expanding walk-forward is exactly the Pareto frontier of causal validation, $k$-fold cross-validation buys the most future data, and purged $k$-fold with an embargo pays in distance instead, which is cheap when the process forgets quickly but cannot repair the part of causality demanded by non-stationarity. On pure noise, shuffled 5-fold reports an information coefficient of $+0.32$, while contiguous 5-fold, using the same amount of future data, reports $+0.004$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑