进化组合模型:面向多目标大语言模型对齐的进化型专家混合模型
Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment
浏览论文内容
中文总结 AI 辅助
针对大语言模型多目标对齐需求,提出进化型专家混合框架,通过进化算法训练门控网络,在多任务上实现了优于基线的超体积、线性效用和切比雪夫效用,提升约20%。
中文摘要 AI 辅助
大型语言模型越来越需要生成满足多个相互冲突目标的响应。由于最优权衡取决于用户偏好和输入提示,可控多目标生成必须在推理时动态适配模型,且无需重新训练。为解决该问题,我们提出进化型专家混合模型(Evolutionary Soups),这是一个用于细粒度生成控制的专家混合框架,其门控网络通过进化算法进行训练。每层门控网络从隐藏状态表示中动态生成专家合并系数,而进化算法结合贪心超体积贡献,以有效进化这些门控网络,在大规模且含噪声的训练数据集上实现持续改进,并覆盖非凸帕累托前沿的更广泛范围。在三个任务上开展的实验证明了Evolutionary Soups相比基线方法的有效性:在所有任务的可控方法中,它取得了最佳的超体积、线性效用和切比雪夫效用(提升约20%)。
英文摘要
Large language models are increasingly required to generate responses that satisfy multiple competing objectives. Since optimal trade-offs depend on both user preferences and input prompts, controllable multi-objective generation must dynamically adapt models at inference time without retraining. To address this, we propose Evolutionary Soups, a mixture-of-experts framework for fine-grained generation control, with gating networks trained via an evolutionary algorithm. The per-layer gating networks dynamically produce expert-merging coefficients from hidden-state representations, while the evolutionary algorithm incorporates greedy hypervolume contribution for effective evolution of these gating networks, achieving consistent improvements on large and noisy training datasets and broader coverage of the non-convex Pareto front. Experiments across three tasks demonstrate the effectiveness of Evolutionary Soups over baselines: it achieves the best hypervolume, linear utility, and Tchebyshev utility (~20% improvement) among controllable methods on all tasks.
发表机构
- Fraunhofer Institute for Applied Information Technology FIT(弗劳恩霍夫应用信息技术研究所FIT)
- University of Cologne(科隆大学)
- University of Stuttgart(斯图加特大学)
- University of Southampton(南安普顿大学)
- Soochow University(苏州大学)
- University Hospital of Cologne(科隆大学医院)
机构由 AI 辅助整理,请以论文原文为准。