arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个补丁,三个角色:自回归时间序列预测中究竟耦合了什么?

One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?

Ziang Li, Yue Huang, Guoxu Zhou, Na Han, Jie Wen, Lunke Fei, Xiaozhao Fang

arXiv 2609.23686首次发表:更新:

发表机构

Guangdong University of Technology; Guangdong Polytechnic Normal University; Harbin Institute of Technology(广东工业大学; 广东技术师范大学; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究解耦自回归时间序列预测中补丁长度对输入表示、转移和执行的耦合,提出轨迹蒸馏与频谱切线校正,实现加速与误差降低。

AI 中文摘要

基于补丁的自回归时间序列预测通常将输入表示、学习到的转移和递归执行绑定到一个补丁长度上。我们探究这些角色中哪些可以单独调整。一项支持性的原子编码研究发现,在评估的网格上,模型宽度比原子分组更敏感。我们的主要发现是,冻结的父模型的递归轨迹比通过轻量级并行出口观察到的未来更容易拟合。自回归轨迹蒸馏(ATD)将其转化为可选择的ATD-1/2/4/8执行,其中ATD-1完全恢复父模型。在配对的四数据集比较中,ATD-8实现了5.54倍的端到端加速,且在不同宽度下质量稳定。较少的调用并不会自动消除父模型现有的预测误差:在所有21次种子运行中,ATD提高了轨迹保真度,但在与匹配的干净未来监督相比时,仅在15次中提高了预测准确性。我们进一步发现了一个可纠正的残差投影,沿着训练选择的周期性历史方向。频谱切线在不增加神经参数或Transformer调用的情况下应用此校正。在水平线720处,它在七个数据集和两个输出宽度上将均方误差(MSE)和平均绝对误差(MAE)分别降低了2.54%和2.33%,同时比递归推理快3.24倍。水平和形状投影有时不一致。轨迹可压缩性、保真度-准确性不匹配以及校正现象在三个公共AR父模型中重复出现。这些结果共同将表示、转移和执行分离为AR设计轴。代码可在https://this URL获取。

英文摘要

Patch-based autoregressive time-series forecasting often ties input representation, learned transitions, and recursive execution to one patch length. We ask which of these roles can be adjusted separately. A supporting atomic-encoding study finds greater sensitivity to model width than to atom grouping on the evaluated grid. Our main finding is that a frozen parent's recursive trajectory is easier to fit than the observed future with lightweight parallel exits. Autoregressive Trajectory Distillation (ATD) turns this into selectable ATD-1/2/4/8 execution, with ATD-1 exactly recovering the parent. On a paired four-data-set comparison, ATD-8 reaches $5.54\times$ end-to-end speedup with stable quality across widths. Fewer calls do not automatically remove the parent's existing forecast error: ATD improves trajectory fidelity in all 21 seed runs but forecast accuracy in only 15 against matched clean-future supervision. We further find a correctable residual projection along a train-selected periodic history direction. Spectrum Tangent applies this correction without adding neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% over seven data sets and two output widths, while remaining $3.24\times$ faster than recursive inference. Level and shape projections sometimes disagree. Trajectory compressibility, the fidelity-accuracy mismatch, and the correction recur across three public AR parents. Together these results separate representation, transition, and execution as AR design axes. Code is available at https://github.com/RowanFFF/ATD-Spectrum-Tangent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑