arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06421stat.ML

不平衡学习优化器比较研究

On the Comparison of Optimizers for Imbalanced Learning

Cl{é}ment Lezane, Fran{\c c}ois Bachoc, J{é}r{ô}me Bolte, Jean-Michel Loubes

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究数据不平衡下的优化器比较,通过连续时间几何建模刻画多数类优化区域及离开时间界限,发现符号、谱和牛顿下降对少数类振幅依赖更温和,实验验证AdamW和Muon优于SGD。

中文摘要 AI 辅助

数据不平衡在机器学习中普遍存在,从稀有词汇和异常值到异构或交叉表格数据中代表性不足的模式。我们研究连续时间下的理想化优化器几何结构,以模拟深度学习中的小步长训练。我们假设不平衡的来源未被观测到:优化器只能访问聚合训练损失,而忽略多数类和少数类的确切贡献。在此设定下,我们刻画了一个区域,在该区域中,无论少数类的可容许结构如何,多数类损失都会被优化。我们推导出该区域的显式方程以及离开该区域所需时间的界限。这些界限表明,对于符号、谱和牛顿下降法,其对少数类振幅的依赖性比欧几里得梯度下降法更温和。使用AdamW和Muon进行的实验表明,在语言、表格和图像任务上,它们相对于SGD具有类似优势。

英文摘要

Data imbalance is pervasive in machine learning, from rare words and anomalies to underrepresented patterns in heterogeneous or cross-tabulated data. We study idealized optimizers geometries in continuous time to model small-step training in deep learning. We assume that the source of imbalance is unobserved: the optimizer has only access to the aggregate training loss ignoring the exact contributions of the majority and minority groups. In this setting, we characterize a region where majority losses are optimized regardless of the admissible minority structure. We derive explicit equations of this zone and bounds on the time needed to leave it. These bounds exhibit a milder dependence on minority amplitude for sign, spectral, and Newton descent than for Euclidean gradient descent. Experiments with AdamW and Muon suggest similar advantages over SGD across language, tabular, and image tasks.

发表机构

  • Université de Toulouse(图卢兹大学)
  • University of Lille(里尔大学)
  • Toulouse School of Economics(图卢兹经济学院)
  • Institut Universitaire de France (IUF)(法国高等研究院)
  • ANITI
  • INRIA(法国国家信息与自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑