非单调直接搜索方法用于确定性和随机无导数优化
Non-monotone direct-search methods for deterministic and stochastic derivative-free optimization
浏览论文内容
中文总结 AI 辅助
本文提出最大-M非单调直接搜索方法,用于确定性和随机无导数优化,通过允许目标函数暂时增加来避免次优解,并建立了与单调方法匹配的O(ε^{-2})迭代复杂性理论。
中文摘要 AI 辅助
在无导数优化(DFO)中,人们最小化梯度不可用或计算代价高昂的函数。在许多应用中,由于模拟或系统随机性,目标函数值和梯度带有噪声。一类标准的直接搜索方法在试验点使目标函数减少量与步长平方成比例时接受该试验点。然而,当应用于复杂地形时,这种要求可能使算法陷入次优解的邻域。我们研究了一种非单调直接搜索替代方案,其中试验函数值与最近 $M$ 个不同迭代点中获得的最大目标函数值进行比较。这种最大-$M$ 非单调条件允许目标函数暂时增加,并有助于穿越狭窄的弯曲山谷;然而,由于缺乏单调递减,其理论分析显著更具挑战性。在本文中,我们为确定性和随机DFO问题中的最大-$M$ 非单调直接搜索建立了全面的复杂性理论。对于确定性目标,我们基于正生成集建立了完全轮询的最坏情况迭代界,以及基于概率下降轮询的期望迭代界。然后,我们分析了使用独立函数估计的随机变体,并在随机误差的尾界假设下展示了期望迭代复杂性。所有三个结果都具有标准的 $\nmathcal{O}(\nepsilon^{-2})$ 复杂性,这与单调直接搜索方法的迭代复杂性相匹配。我们的理论由一组新的 merit 函数实现,这些函数通过步长平方的有序倍数校正存储的目标值,并结合概率方法的更新-奖励停止时间论证。
英文摘要
In derivative-free optimization (DFO), one minimizes functions for which the gradient is unavailable or expensive to compute. In many applications, objective function values and gradients are noisy due to simulations or system randomness. A class of standard direct-search methods for DFO accept a trial point when it decreases the objective function by an amount proportional to the squared stepsize. However, when applied to complex landscapes, such a requirement may trap the algorithm in a neighborhood of sub-optimal solutions. We study a non-monotone direct-search alternative where the trial function value is compared with the largest objective function obtained through the $M$ most recent distinct iterates. This max-$M$ non-monotone condition permits temporary increases in the objective function and can help navigate narrow curved valleys; however, its theoretical analysis is significantly more challenging due to the lack of monotonic decrease. In this paper, we develop a comprehensive complexity theory for the max-$M$ non-monotone direct-search in both deterministic and stochastic DFO problems. For deterministic objectives, we establish a worst-case iteration bound for a complete poll based on a positive spanning set and an expected iteration bound for a probabilistic-descent poll. We then analyze a stochastic variant using independent function estimates and show the expected iteration complexity under tail-bound assumptions of the stochastic errors. All three results have the standard complexity of $\mathcal{O}(ε^{-2})$, which matches the iteration complexity of monotone direct-search methods. Our theory is enabled by a new family of merit functions that correct the stored objective values by ordered multiples of the squared stepsize, together with a renewal-reward stopping-time argument for the probabilistic methods.
发表机构
- Lehigh University(利哈伊大学)
- Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。