面向专家混合模型的结构感知神经架构搜索
Structure Aware Neural Architecture Search for Mixture of Experts
浏览论文内容
中文总结 AI 辅助
该研究提出结构感知NAS框架,将数据簇与专家的对齐作为显式搜索变量,通过广义EM流程优化,在无标签的图像分类和时间序列预测任务上优于MoE与NAS基线。
中文摘要 AI 辅助
神经架构搜索(NAS)迄今很少应用于专家混合(MoE)模型,现有的MoE设计让专家与数据结构的对齐过程自主产生。我们提出一种架构搜索框架,将这种对齐设为显式搜索变量:数据簇到专家的分配与每个专家的架构同步优化。我们将该联合问题转化为感知簇的似然最大化,证明其与隐变量混合模型的不完整数据最大似然等价,并通过广义期望-最大化(EM)流程求解,其中原本难以处理的专家质量项由自适应优化的代理提供。我们证明,当代理误差可求和时迭代会收敛,且在每个极限点,搜索产生的候选解不会提升真实目标。在异构图像分类混合任务上,该方法在未观察到域标签的情况下,对95%的簇恢复了潜在域划分;在该基准任务和四域时间序列预测任务上,它均优于同样未使用标签信息的MoE和NAS基线方法。
英文摘要
Neural Architecture Search (NAS) has so far rarely been applied to Mixture-of-Experts (MoE) models, and existing MoE designs leave the alignment between experts and the structure of the data to emerge on its own. We propose an architecture search framework that makes this alignment an explicit search variable: the assignment of data clusters to experts is optimised jointly with the per-expert architectures. We cast the joint problem as a cluster-aware likelihood maximisation, show that it coincides with the incomplete-data maximum likelihood of a latent-variable mixture, and solve it by a generalised Expectation-Maximisation procedure whose otherwise intractable expert-quality term is supplied by an adaptively refined surrogate. We prove that the iterates converge whenever the surrogate errors are summable, and that at every limit point no candidate the search produces improves the true objective. On a heterogeneous image-classification mixture the method recovers the underlying domain partition on 95% of clusters without ever observing domain labels, and on that benchmark and a four-domain time-series forecasting one alike it outperforms the MoE and NAS baselines that likewise use no label information.