发表机构
University of Science and Technology Beijing; Zhejiang University of Technology; Institution of Artificial Intelligence, University of Science and Technology Beijing(北京科技大学; 浙江工业大学; 北京科技大学人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SACQ是一种插件式结构化预测头,通过两阶段解码流程和批量自适应缩放对数-cosh损失,在PatchTST等骨干网络上实现长时序预测的顶级性能,且对输入损坏和标签噪声的鲁棒性显著优于扁平化读出头。
AI 中文摘要
长期时序预测(LTSF)模型主要采用基于patch的编码器,其末端为扁平化读出头,通过单个共享投影将整个编码后的历史记忆映射到所有未来步。这种未来位置的隐式耦合会掩盖位置特定的历史-未来对齐关系,并放大对损坏输入和极端监督噪声的敏感性。我们提出SACQ,这是一种插件式结构化预测头,可在保持编码器不变的情况下替换扁平化读出头。SACQ采用两阶段解码流程:首先建立粗略的patch网格预测支架,然后通过对历史记忆的交叉注意力优化每个未来位置,并通过学习到的每个patch门将注意力衍生的修正与粗略支架合并。为了在长时序和噪声标签下稳定优化,我们进一步提出批量自适应缩放对数-cosh损失,该损失可自动校准对当前残差尺度的鲁棒性,抑制异常梯度的同时保留类似MSE对典型误差的敏感性。SACQ在PatchTST、DLinear和patch-Mamba骨干网络上达到了顶级的测试MSE/MAE,且仅带来适度的参数和延迟增量开销。在推理时输入损坏和训练集标签噪声的压力测试中,SACQ显著优于扁平化读出头, ablation研究验证了每个架构组件的有效性。
英文摘要
Long-term time series forecasting (LTSF) models predominantly employ patch-based encoders terminated by a flatten readout head that maps the entire encoded historical memory to all future steps through a single shared projection. This implicit coupling of future positions obscures position-specific historical-to-future alignment and amplifies sensitivity to corrupted inputs and extreme supervision noise. We present SACQ, a plug-in structured prediction head that replaces flatten readout while keeping the encoder unchanged. SACQ adopts a two-stage decoding pipeline: it first establishes a coarse patch-grid forecast scaffold, then refines each future position through cross-attention over historical memory and merges the attention-derived correction with the coarse scaffold via a learned per-patch gate. To stabilize optimization under long horizons and noisy labels, we further propose a batch-adaptive scaled log-cosh loss that automatically calibrates robustness to the current residual scale, suppressing outlier gradients while preserving MSE-like sensitivity for typical errors. SACQ attains top-tier test MSE/MAE across PatchTST, DLinear, and patch-Mamba backbones with only modest incremental overhead in parameters and latency. Under inference-time input corruption and training-set label-noise stress tests, SACQ substantially outperforms flatten readouts, with ablation studies validating each architectural component.
Comments10 pages, 7 figures