arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有变量一致:面向多元时间序列预测的可靠性感知变量级梯度手术

Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting

Jinwoo Park, Hyeongwon Kang, Pilsung Kang

arXiv 2609.08554首次发表:更新:

发表机构

Seoul National University; Korea University(首尔大学; 高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多元时间序列预测中均值损失掩盖变量间梯度冲突的问题,提出可靠性感知的逐变量梯度手术方法PV-Surgery,通过代理梯度与条件池化优化,平均降低MSE 3.61%和MAE 2.93%。

AI 中文摘要

在数据驱动的训练中,多元时间序列预测通常使用在样本、变量和预测时域上平均的标量损失进行优化。这种平均虽然方便,但优化器只能看到聚合梯度,无法揭示变量级贡献是相互一致还是相互对立。为了量化这种不一致出现的频率,我们直接测量了变量级梯度,发现在七个数据集上,其两两余弦相似度平均有30.6%为负值。然而,冲突与损害并非同一回事。在共享训练下,64个变量中有35个的表现不如全输入单目标基准模型,且受损比例并不能被梯度冲突频率可靠预测。我们提出了逐变量手术(PV-Surgery),一种面向具有缓存兼容层的主干网络的优化器端训练策略。一次反向传播利用输出端信号构建变量级梯度代理,并保留逐点预测损失。可靠性感知选择针对代理和与其共享梯度切片近似接近的层。条件池化在不丢弃变量的情况下形成锚定池和冲突池。公共方向手术将变量或池化梯度与其归一化均值对齐,并恢复输入范数以避免重新加权。在五个主干网络、七个数据集和四个预测时域上的实验中,PV-Surgery平均降低了MSE 3.61%和MAE 2.93%。对于多元预测,这表明被均值损失训练所隐藏的变量级结构是一种可用的优化信号。

英文摘要

In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another. To quantify how often this disagreement arises, we measure the variable-wise gradients directly and find that 30.6% of their pairwise cosine similarities are negative on average across seven datasets. However, conflict and harm are not the same thing. Under shared training 35 of the 64 variables do worse than a full-input single-target oracle, and the harmed fraction is not reliably predicted by how often gradients conflict. We propose Per-Variable Surgery (PV-Surgery), an optimizer-side training strategy for backbones with cache-compatible layers. One backward pass builds variable-wise gradient proxies from output-side signals and keeps the pointwise forecasting loss. Reliability-aware selection targets layers whose proxy sums closely approximate their shared-gradient slices. Conditional pooling forms anchor and conflict pools without dropping variables. Common-direction surgery aligns variable or pooled gradients with their normalized mean and restores input norms to avoid reweighting. In experiments across five backbones, seven datasets, and four horizons, PV-Surgery lowers MSE by 3.61% and MAE by 2.93% on average. For multivariate forecasting, this indicates that the variable-wise structure hidden by mean-loss training is a usable optimization signal.

Comments34 pages, 21 figures, 20 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑