arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于学习优化的高效长时学习

Efficient Long-Horizon Learning for Learned Optimization

Xiaolong Huang, Benjamin Thérien, James Harrison, Eugene Belilovsky

arXiv 2607.06772首次发表:更新:

发表机构

Mila - Quebec AI Institute; Google DeepMind; Concordia University; Université de Montréal(米拉-魁北克人工智能研究所; 谷歌DeepMind; 康考迪亚大学; 蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对学习优化中当前元训练方法的局限,提出高效长时(ELO)学习算法,重新分配计算并实施监督,提升长展开性能和分布外泛化能力,在多任务中表现出色,且元训练所需GPU时长少。

AI 中文摘要

学习优化旨在通过在任务分布上进行元学习小型神经网络优化器来改进手工设计的优化器(如Adam和Muon)。近期工作虽推进了学习优化器(LOs)的架构设计和归纳偏差,但当前元训练方法仍有两个主要困难:无法有效扩展到长时内部问题,且常无法超越手工设计的优化器。为解决这些局限,我们提出高效长时(ELO)学习,它重新分配冗余元训练计算到更长失败阶段以实现高效长时学习,还实施解耦渐进专家监督以提供稳定元学习信号并提升LOs泛化能力。实证研究评估了ELO在按元素和基于矩阵的LOs元训练中的效果。在下游语言建模和图像分类任务中,ELO显著提升了基础LOs的长展开性能和分布外泛化能力。特别是ELO - Celo2在所有评估任务中持续优于调优良好的AdamW,在语言建模上与Muon竞争。值得注意的是,所有ELO基线在元训练时所需的H100 GPU时长不到7小时。

英文摘要

Learned optimization aims to improve upon hand-designed optimizers (e.g., Adam and Muon) by meta-learning small neural network optimizers over a distribution of tasks. While recent work has greatly advanced the architectural design and inductive biases of learned optimizers (LOs), their meta-training remains biased toward short-unroll learning on particular tasks, resulting in redundant computation and leaving LOs often unable to compete with hand-designed optimizers. We introduce Efficient Long-hOrizon (ELO) learning, an efficient meta-training algorithm that (1) reallocates wasted meta-training compute to longer failure regimes, achieving efficient long-horizon learning, and (2) enforces decoupled progressive expert supervision, providing stable meta-learning signals that additionally improve the generalization of LOs. Our empirical study evaluates ELO for meta-training both element-wise and matrix-based LOs. Across downstream language modeling (GPT-2-124M/350M on FineWeb) and image classification (ViT-B/16, ResNet-50 on ImageNet-1K) tasks, ELO substantially improves the long-unroll performance and out-of-distribution generalization of the base LOs. In particular, ELO-Celo2 consistently outperforms well-tuned AdamW across all evaluated tasks, while remaining competitive with Muon on language modeling. \textit{Notably, all ELO baselines require less than 7 H100 GPU-hours for meta-training.}

CommentsMeta-learning, learned optimization

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑