更多专家,更差的动力学:混合专家状态空间模型中的逆缩放与谱偏置
More Experts, Worse Dynamics: Inverse Scaling and Spectral Bias in Mixture-of-Experts State-Space Models
浏览论文内容
中文总结 AI 辅助
该研究在合成环境中评估混合专家状态空间模型,发现增加专家数量会导致逆缩放等问题,算子级混合模型未优于单专家基线,需几何感知评估模式切换动力学系统。
中文摘要 AI 辅助
混合专家(MoE)架构通常被认为是通过将复杂系统分解为更简单的局部动力学来提升表达能力的一种方式。最近,这种直觉被扩展到谱状态空间模型中,在该模型中,假设混合稳定算子能够适应异构或存在模式切换的时间序列。我们在旨在隔离动力学而非表征挑战的受控合成环境中批判性地评估这一假设。我们研究由三种模式组成的序列的下一步预测任务:由Mackey-Glass系统生成的混沌动力学、稳定振荡模式以及噪声主导的自回归模式。在包括容量缩放、神谕路由、冻结专家变体以及与输出级MoE基线的比较在内的大量消融实验中,算子级混合模型始终未能优于单专家基线。增加专家数量会导致逆缩放、路由崩溃或无法诱导有意义的专业化,即使是完美的模式监督也无法防止全局性能下降。此外,我们表明混沌轨迹上均方误差的表观改进可能具有误导性。相空间分析显示,较低的误差通常源于破坏底层吸引子几何结构的时间平滑,而非对动力学的忠实建模。这些结果确定了所研究参数化和训练协议下算子插值的可能局限,并强调在评估模式切换动力学系统时需要几何感知的评估。
英文摘要
Mixture-of-Experts (MoE) architectures are commonly motivated as a way to increase expressivity by decomposing complex systems into simpler local dynamics. This intuition has recently been extended to spectral state-space models, where mixing stable operators is assumed to enable adaptation to heterogeneous or regime-switching time series. We critically evaluate this assumption in a controlled synthetic setting designed to isolate dynamical rather than representational challenges. We study a next-step prediction task on sequences composed of three regimes: chaotic dynamics generated by the Mackey-Glass system, a stable oscillatory regime, and a noise-dominated autoregressive regime. Across extensive ablations including capacity scaling, oracle routing, frozen-expert variants, and comparisons to output-level MoE baselines, operator-level mixture models consistently fail to outperform a single-expert baseline. Increasing the number of experts leads to inverse scaling, routing collapses or fails to induce meaningful specialization, and even perfect regime supervision does not prevent degradation in global performance. Furthermore, we show that apparent improvements in mean squared error on chaotic trajectories can be misleading. Phase-space analysis reveals that lower error often arises from temporal smoothing that destroys the geometry of the underlying attractor rather than from faithful modeling of the dynamics. These results identify a likely limitation of operator interpolation under the studied parameterization and training protocol, and underscore the need for geometry-aware evaluation when assessing regime-switching dynamical systems.
发表机构
- Delhi Technological University(德里理工大学)
机构由 AI 辅助整理,请以论文原文为准。