AI 中文总结
本研究针对非线性非平稳系统辨识,提出上下文感知端到端MoE框架,在NARMA-10任务上较基线模型大幅降低工况切换后的测试误差,提升鲁棒性。
AI 中文摘要
本研究针对非线性非平稳系统辨识问题,采用具有十步记忆的非线性自回归滑动平均基准(NARMA-10)开展一步 ahead 预测。基线模型包括带 exogenous 输入的自回归模型(ARX)、采用多层感知机的非线性自回归模型(NARX-MLP)及单个长短期记忆网络(LSTM),用于对比预测性能。为提升鲁棒性,本研究通过凸混合方式整合多个 LSTM 专家。标准自适应凸混合采用基于误差驱动的规则结合单纯形投影更新混合权重,但该机制具有反应性、需手动调参且无法端到端学习。所提方法引入上下文感知的端到端专家混合(MoE)框架,其中可微 softmax 门控网络与专家参数联合学习上下文感知混合权重。在平稳 NARMA-10 任务上,所提 MoE 门控方法性能与自适应凸混合相当;在工况切换动态下,受控冻结专家的 ablation 实验分离了混合权重更新机制,结果显示 MoE 门控显著提升鲁棒性,整体测试误差降低约五倍,切换后误差降低约二点五倍。
英文摘要
This work addresses nonlinear and nonstationary system identification using one-step-ahead prediction on the nonlinear autoregressive moving average benchmark with ten-step memory (NARMA-10). Baseline models, including autoregressive models with exogenous input (ARX), nonlinear ARX using a multilayer perceptron (NARX--MLP), and a single long short-term memory network (LSTM), are used to contextualize prediction performance. To improve robustness, multiple LSTM experts are combined through a convex mixture. A standard adaptive convex mixture updates the mixing weights using an error-driven rule with simplex projection, but this mechanism is reactive, hand-tuned, and not end-to-end learnable. The proposed method introduces a context-aware end-to-end mixture-of-experts (MoE) framework in which a differentiable softmax gating network learns context-aware mixing weights jointly with the expert parameters. On stationary NARMA-10, the proposed MoE gating approach achieves performance comparable to the adaptive convex mixture. Under regime-switching dynamics, a controlled frozen-experts ablation isolates the mixing-weight update mechanism and shows that the MoE gate significantly improves robustness, achieving approximately fivefold lower overall test error and about two-and-a-half-fold lower after-switch error.
Comments6 pages, 10 figures, 3 tables