arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29774cs.LG

辅助学习的解析理论

An Analytical Theory of Auxiliary Learning

Federico Milanesio, Alessandro Ingrosso, Matteo Osella

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过教师-学生框架和微分方程解析了辅助学习机制,推导出线性网络泛化误差闭式解及非线性情形的涨落-耗散理论,揭示任务相关性与噪声的作用。

中文摘要 AI 辅助

辅助学习是一种优化范式,其中通过在额外任务上联合训练神经网络,可以提高其在目标任务上的性能。然而,这种改进背后的机制仍鲜为人知。我们使用教师-学生框架研究该问题,并在大输入极限下推导出描述在线随机梯度下降动力学的封闭微分方程组。对于线性网络,我们获得了泛化误差在学习率前导阶上的闭式表达式,量化了任务相关性和标签噪声如何决定辅助学习的收益。对于非线性激活函数,我们发展了一种涨落-耗散解析理论,建立了将主任务误差和辅助任务误差与相应的单任务误差联系起来的一般关系。数值实验支持理论预测,并展示了辅助任务如何通过平衡朝向最优解的强迫动力学与梯度噪声来改善泛化。

英文摘要

Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.

发表机构

  • University of Turin(都灵大学)
  • Radboud University(拉德堡德大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑