多智能体系统的动态专家剪枝
Dynamic Expert Pruning for Multi-Agent Systems
浏览论文内容
中文总结 AI 辅助
针对多智能体系统中静态专家剪枝无法适应异构工作负载的问题,提出动态专家剪枝(DEP),利用提示文本预测每请求专家掩码,在多种任务和模型上优于静态基线,且专家越少优势越大。
中文摘要 AI 辅助
混合专家(Mixture-of-Experts, MoE)架构通过仅为每个令牌激活少数专家来高效扩展语言模型,但节省仅限于计算:每个专家必须常驻于加速器上,因此内存限制了这些模型的部署位置。专家剪枝可减少这一内存占用,然而现有方法是静态的——在离线状态下校准的单一掩码被应用于模型的每一个后续请求。当工作负载异构时,这一假设可能失效,最显著的是在多智能体系统中,一个骨干模型同时服务于多个任务和角色:我们的分析表明,不同任务和角色会调用不同的专家,而静态方法为所有任务和角色分配一个固定的子集。因此,我们提出了动态专家剪枝(Dynamic Expert Pruning, DEP),其基于我们在此确立的一个发现:智能体的系统提示和任务提示本身足以识别该智能体及其任务所需的专家,因为该文本已描述了智能体将要执行的操作。一个轻量级预测器,在工作流转录上训练一次,通过单次前向传播将这些提示转换为专门针对每个请求的掩码,无需针对每个配置进行校准。在多样化的任务和角色、模型规模以及MoE架构中,DEP相比静态剪枝和合并基线实现了更好的整体准确率,并且无需重新训练即可泛化到训练中未见的工作流。其相对于这些基线的优势在保留专家数量较少时最大,这表明多智能体系统固有的角色专业化允许比静态剪枝更稀疏的服务。
英文摘要
Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning reduces this footprint, yet existing methods are static --- a single mask, calibrated offline, is applied to the model for every subsequent request. This assumption can fail when the workload is heterogeneous, most prominently in multi-agent systems, where one backbone serves many tasks and roles at once: our analysis shows that different tasks and roles recruit different experts, while static methods assign one fixed subset to all of them. We therefore propose Dynamic Expert Pruning (DEP), which rests on a finding we establish here: an agent's system and task prompts are by themselves sufficient to identify the experts that agent and its task require, since that text already describes what the agent will do. A lightweight predictor, trained once on workflow transcripts, turns those prompts into a specialized per-request mask in a single forward pass, with no per-configuration calibration. Across diverse tasks and roles, model scales, and MoE architectures, DEP achieves better overall accuracy than static pruning and merging baselines, and generalizes to workflows unseen in training without retraining. Its margin over those baselines is largest when few experts are retained, suggesting that the role specialization inherent to multi-agent systems permits sparser serving than static pruning allows.
发表机构
- Pohang University of Science and Technology (POSTECH)(浦项科技大学)
- Microsoft Research Asia(微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。