arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

原始路由混合适配器:时间序列基础模型中路由崩溃的因果干预

Raw-Routed Mixture of Adapters: A Causal Intervention for Routing Collapse in Time Series Foundation Models

Hung Phan, Thuy T. Nguyen, Minh Ngoc Dinh, Nhat-Quang Tran

arXiv 2609.39445首次发表:更新:

AI 中文总结

针对时间序列基础模型在实例归一化骨干上路由崩溃的问题,提出原始路由混合适配器(RR-MoA),通过因果干预在归一化前输入上路由,在冻结骨干下显著优于多种适配方法,并具有跨骨干泛化性。

AI 中文摘要

时间序列基础模型(TSFMs)通常通过将单个可训练头部附加到冻结骨干上来适应新数据,这是一种一刀切的设置,无法充分拟合异构机制。将头部替换为专家混合是标准的升级方式,但在实例归一化骨干(主导的TSFM设计类别)上,该方式失败:路由熵降至零,一个专家吸收所有输入,我们将这种失败称为归一化诱导的路由崩溃。标准的MoE救援机制无法修复此问题,因为原因在于路由器的输入,而非其优化过程。编码器前的归一化剥离了路由器区分不同机制所需的统计信息。互信息分解使这一点更加精确,并产生一个信号比率,该比率在训练前计算,可预测数据集脆弱性(Spearman ρ = -0.88)。包括视觉模态复制在内的八项因果对照实验,将实例归一化确定为原因。解决方案是一种最小因果干预:原始路由混合适配器(RR-MoA),它在原始、归一化前的输入上进行路由。在严格冻结骨干下,RR-MoA在54/54次比较中胜过最强的固定适配器,并显著优于LoRA、TRACE、AdaMix和全参数微调。该效果在六个骨干和一个插补任务上具有泛化性。冻结的RR-MoA还比全参数微调高出12-79%(冻结悖论);两个架构不同的变体证实了该原理超越了这一特定路由器。

英文摘要

Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup that underfits heterogeneous regimes. Replacing the head with a mixture of experts is the standard upgrade, but on instance-normalized backbones (the dominant TSFM design class) it fails: routing entropy collapses to zero and one expert absorbs every input, a failure we call normalization-induced routing collapse. Standard MoE rescue mechanisms do not repair it, because the cause is in the router's input, not its optimization. Pre-encoder normalization strips the statistics a router would need to tell regimes apart. A mutual-information decomposition makes this precise and yields a signal-ratio that, computed before training, predicts dataset vulnerability (Spearman $ρ= -0.88$). Eight causal controls, including a vision-modality replication, isolate instance normalization as the cause. The prescription is a minimal causal intervention: Raw-Routed Mixture of Adapters (RR-MoA), which routes on the raw, pre-normalization input. Under a strictly frozen backbone, RR-MoA wins 54/54 comparisons against the strongest fixed adapter and significantly outperforms LoRA, TRACE, AdaMix, and full fine-tuning. The effect generalizes across six backbones and an imputation task. Frozen RR-MoA also beats full fine-tuning by 12-79% (the Frozen Paradox); two architecturally distinct variants confirm the principle generalizes beyond this specific router.

CommentsAccepted at NeurIPS 2026 (poster). 50 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑