arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度学习模型组合的分布优化视角

A Distributional Optimisation Perspective on Combining Models in Deep Learning

Congye Wang, Yan Lin, Zheyang Shen, Matthew A. Fisher, Chris. J. Oates

arXiv 2609.24328首次发表:更新:

发表机构

Newcastle University(纽卡斯尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从分布优化视角统一集成学习和低秩适配器平均,证明前者凸性保证收敛,并评估新算法,在合成分类和LLM微调上验证有效性。

AI 中文摘要

将不同模型的预测结果进行组合可以提升机器学习任务的性能,但各个模型的训练以及用于组合它们的规则通常是分别选择的,并且往往采用临时性的手段。分布优化(即在概率分布集合上进行优化)领域的最新进展为有原则的联合训练提供了机会,将模型集合视为一个离散分布,其支撑点有待优化,但这些方法的潜力尚未被充分理解。在本文中,我们(1)将两种标准的组合策略——集成学习和低秩适配器平均——表述为熵正则化的分布优化,观察到由此产生的目标函数在集成情形下是凸的,但在适配器平均情形下不是凸的,因此均值场朗之万动力学的现有收敛保证仅适用于前者;(2)评估了针对此任务的现有算法和新算法,包括变分梯度下降的一种函数变体;(3)报告了一项实证研究,涵盖合成分类任务以及在常识推理基准上对大型语言模型的微调。

英文摘要

Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set of probability distributions) offer an opportunity for principled joint training, viewing the collection of models as a discrete distribution whose support points are to be optimised, but the potential of these methods is not well-understood. In this paper we (1) cast two standard combination strategies - ensembles and low-rank adapter averaging - as entropy-regularised distributional optimisation, observing that the resulting objective is convex in the ensemble case but not in the adapter-averaging case, so that existing convergence guarantees for mean field Langevin dynamics transfer only to the former; (2) assess existing and novel algorithms for this task, including a functional variant of variational gradient descent; and (3) report an empirical study spanning synthetic classification tasks and fine-tuning of large language models on a commonsense reasoning benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑