arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文路由何时有用?时间序列预测中多模态融合的系统研究

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin, Yixuan Shen

arXiv 2608.25128首次发表:更新:

发表机构

University of Tennessee, Knoxville; WorkMagic; University of Texas at Austin(田纳西大学诺克斯维尔分校; WorkMagic; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究明确了多模态时间序列预测中辅助上下文有效的两个数据集条件,通过实验证实仅当两条件满足时,文本条件专家调制可显著降低均方误差,且确立了因果关系。

AI 中文摘要

多模态时间序列预测方法通过日益复杂的融合机制将辅助上下文整合到时间预测中。越来越多的研究报告了显著的收益,但通常不清楚这些收益是否反映了对上下文的真正利用,还是偶然的架构效应。我们提出一个更具体、可验证的问题:辅助上下文究竟何时能帮助预测器?我们确定了两个必须同时满足的数据集级条件:(1) 目标不受最后值捷径主导(低自相关系数ρ_h);(2) 上下文携带超出历史的目标信息(非零条件互信息δ;当δ=0时,任何预测器都无法从中受益——这是一个与分布无关的结果)。通过在MoME(一个143亿参数的混合专家模型,6个数据集,10个随机种子)和在单主干测试平台中实现的另外四种融合机制(5个数据集)上进行受控实验,我们发现当两个条件都满足时,文本条件专家调制会带来可观的均方误差(MSE)降低;当任一条件不满足时,其贡献会降至调制路径的容量下限,且不包含可归因于上下文的信号。我们通过两种干预措施确立因果关系:向MoME添加捷径会在3个数据集上抑制77%-93%的路由贡献;逐步破坏上下文质量会使上下文特定收益从+44%变为负值。我们在27个Monash Archive数据集上验证了诊断方法的自相关组件。我们提供了一个校准的预训练诊断,在我们测试的数据集上,在足够功效的设置中没有出现假阳性。我们明确说明证据的不对称性:否定结论总体上可靠,而较大的正值来自单一模型家族(MoME),仅在方向上得到测试平台的证实。

英文摘要

Multi-modal time series forecasting methods integrate auxiliary context into temporal predictions through increasingly sophisticated fusion mechanisms. A growing body of work reports substantial gains, yet it is often unclear whether they reflect genuine use of the context or incidental architectural effects. We ask a narrower, checkable question: when can auxiliary context help a forecaster at all? We identify two dataset-level conditions that must both hold: (1) the target is not dominated by a last-value shortcut (low autocorrelation rho_h), and (2) the context carries information about the target beyond history (non-zero conditional mutual information delta; when delta=0 no predictor can benefit---a distribution-free result). Through controlled experiments on MoME (a 14.3B-parameter mixture-of-experts model, 6 datasets, 10 seeds) and four additional fusion mechanisms implemented within a single-backbone testbed (5 datasets), we find that when both conditions hold, text-conditioned expert modulation contributes a sizeable MSE reduction; when either fails, the contribution collapses to the capacity floor of the modulation pathway and carries no context-attributable signal. We establish causality through two interventions: adding a shortcut to MoME suppresses routing contribution by 77-93% across 3 datasets; progressively corrupting context quality drives the context-specific benefit from +44% to negative. We validate the autocorrelation component of our diagnostic on 27 Monash Archive datasets. We provide a calibrated pre-training diagnostic that, on the datasets we test, yields no false positives in well-powered settings. We are explicit about the asymmetry of our evidence: the negative arm is broadly reliable, while the large positive magnitudes come from a single model family (MoME) and are corroborated only in direction by the testbed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑