发表机构
University of Central Florida; Mohammed VI Polytechnic University(中佛罗里达大学; 穆罕默德六世理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型微调中的步长选择难题,提出零阶与一阶混合优化框架ZFO,解耦方向与步长,利用两次额外评估构建局部模型自适应选步,理论保证收敛,实验显示优于固定步长基线。
AI 中文摘要
步长选择在大规模神经网络优化中仍是一个核心挑战:保守的步长会减慢收敛速度,而激进的步长则可能破坏稳定性。我们结合零阶与一阶优化方法(ZFO),提出了一种轻量级框架,将方向选择与步长选择解耦。ZFO使用可信的一阶优化器来确定方向,仅沿这一维子空间进行零阶评估以决定移动多远。利用当前的梯度信息和两次额外的目标函数评估,ZFO实例沿建议方向构建目标函数的局部模型,并在有界搜索区间内选择曲率感知的步长。这产生了一种自适应步长选择机制,其成本低于完整的线搜索。我们提供了理论保证,证明共享样本评估能产生可靠的有限差分曲率估计,所构建的局部模型能在搜索区间内选择接近最优的步长,并且ZFO能收敛到驻点的一个邻域。在评估的设置、语言模型和数据集上,相对于固定步长的一阶基线,ZFO经常能改善优化过程和最终性能,其改进幅度和偏好的局部模型取决于目标函数。我们的代码公开于:此HTTPS链接。
英文摘要
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.
CommentsAccepted to 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Code: https://github.com/nizswan/Zeroth-First-Order-Framework
Journal ref40th Conference on Neural Information Processing Systems (NeurIPS 2026)