arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12890cs.LGcs.AI

大距离梯度未必可靠:面向长时程自回归预测的可靠性加权信用分配

Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting

  • University of Maryland, College Park(马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

Junhao Zhao, David Michael Simberg, Jacob Kang, Colin Connor Kurniawan, Nan Xu

AI总结:

针对长时程自回归预测中远距离梯度可能不可靠的问题,提出Internal-DW方法,通过可靠性加权内部梯度路径,在多个测试平台上降低预测误差5.2%-13.8%,优于现有方法。

AI中文摘要:

在自回归预测中,长预测展开提供了远距离监督,但通过时间反向传播(BPTT)会将这些损失产生的梯度经由许多自回归步骤传递。重复的雅可比乘积可能使远距离梯度主导更新,同时放大可预测信号与不可预测噪声;因此,大的远距离梯度未必携带可靠的学习信号。基于这一观察,我们提出了内部双维纳路由(Internal-DW),一种原则性的仅反向干预方法,它保留完整的前向展开和所有时域损失,同时对内部梯度路径进行可靠性加权。在每个残差块中,我们为恒等路径和非线性路径推导出有界的维纳增益,以在保留可预测学习信号与抑制不可预测变化之间取得平衡,并基于路径级梯度统计和显式噪声模型估计这些增益。在一个已知梯度信噪比(SNR)的受控系统中,我们展示了远距离梯度可能在其信噪比下降时仍然增长,而Internal-DW在恢复可预测梯度信号方面降低了留出误差并改善了预测。在四个历史主导、弱驱动测试平台上,Internal-DW相对于完整BPTT将预测误差降低了5.2%-13.8%,在所有四个平台上均优于梯度裁剪和雅可比正则化,并在三个平台上优于验证集选择的截断BPTT(TBPTT)。它还在这四个测试平台上扩展或保持了拟合的最优训练时域范围。在整个基准套件中,当前的Internal-DW估计器具有明确的适用边界:当可用历史有限或所选采样器无法代表主导的驱动相关变化时,其益处会减少或逆转。结果表明,保留长时程监督并不需要同等信任每个反向贡献。

英文摘要:

In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate the update while amplifying predictable signal and unpredictable innovation together; a large distant gradient therefore need not carry reliable learning signal. Motivated by this, we introduce Internal Dual-Wiener routing (Internal-DW), a backward-only intervention that preserves the full forward rollout and all step losses while reliability-weighting internal gradient routes. At each residual block, we derive bounded Wiener gains for the identity and nonlinear routes that balance preserving predictable learning signal against suppressing unpredictable variation, and estimate them from route-level gradient statistics and an explicit noise model. In a controlled system with known gradient signal-to-noise ratio (SNR), we show that distant gradients can grow even as their SNR falls, and that Internal-DW reduces error in recovering predictable gradient signals and improves forecasting. On four history-dominated, weak-drive testbeds, Internal-DW reduces forecast error by 5.2%-13.8% relative to full BPTT, outperforms gradient clipping and Jacobian regularization on three testbeds, with similar performance on shear flow, and outperforms validation-selected truncated BPTT (TBPTT) on three. It also extends or preserves the fitted optimal training-horizon range across these four testbeds. Across benchmarks, the current Internal-DW estimator has a clear applicability boundary: its benefit diminishes or reverses when usable history is limited or when the selected sampler fails to represent dominant drive-dependent variation. The results show that retaining long-horizon supervision does not require trusting every backward contribution equally.

补充信息

↑