发表机构
Ant International; The Chinese University of Hong Kong(蚂蚁国际; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出统一缩放定律与学习理论,通过五参数定律预测时间序列基础模型容量缩放,并揭示历史信息通过激活提取支持预测的机制。
AI 中文摘要
我们开发了一个统一缩放定律和统一时间序列学习理论,以理解模型容量和历史信息如何支持预测。在不同的回看长度和预测视界上,我们分析了来自21个检查点在23个数据集-频率任务上的18,768个实验单元,涵盖六个领域。我们的实证方法将局部资源关系整合为一个简约的、拟合的五参数定律:容量增益随历史增加而增加,上下文增益随饱和而减少,视界效应作为共同偏移进入。在没有Toto 2.0的情况下拟合,该定律预测其视界平均容量缩放曲线,在输入长度2048和4096时平均绝对百分比误差分别为1.09%和1.50%。为了理解历史如何支持预测,我们的学习理论使用高斯回归分析规则识别和预测能力。我们假设全样本模型通过累积权重中的信息来学习,而冻结的时间序列基础模型(TSFMs)通过激活提取信息来使用历史。匹配历史比较确立了额外历史的预测价值。受控参数交换和激活干预提供了证据,表明历史衍生的规则信息可以被保留、跨查询重用,并用于恢复对长上下文预测的贡献。总之,这些发现为容量缩放、上下文分配以及保留和应用历史规则的模型开发提供了信息。代码和主要结果可在该https URL获取。
英文摘要
We develop a Unified Scaling Law and a Unified Theory of Time Series Learning to understand how model capacity and historical information support forecasting. Across different lookback lengths and forecast horizons, we analyze 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks spanning six domains. Our empirical methodology integrates local resource relations into a parsimonious, fitted five-parameter law: capacity gains increase with history, context gains diminish toward saturation, and horizon effects enter as a common shift. Fitted without Toto 2.0, the law predicts its horizon-averaged capacity-scaling curves with mean absolute percentage errors of 1.09% and 1.50% at input lengths 2048 and 4096. To understand how history supports prediction, our learning theory uses Gaussian regression to analyze rule identification and predictive capability. We hypothesize that full-shot models learn by accumulating information in weights, while frozen time series foundation models (TSFMs) use history by extracting information through activations. Matched-history comparisons establish the predictive value of additional history. Controlled parameter exchanges and activation interventions provide evidence that history-derived rule information can be retained, reused across queries, and used to recover a contribution to long-context prediction. Together, these findings inform capacity scaling, context allocation, and the development of models that retain and apply historical rules. Code and main results are available at https://github.com/Fifthky/UniScale.