AI 中文总结
本文针对随时间演化的在线非凸优化问题,提出随机两点直接搜索算法,推导了不同探测率下的迭代复杂度界,可用于求解动力系统最优控制问题。
AI 中文摘要
零阶(即博弈)反馈下的优化是许多目标和/或约束的解析形式不可用的工程问题的核心。在现代应用如在线控制和在线学习中,优化问题常随时间演化,需要自适应优化方法。然而,现有方法大多局限于为时间不变或一阶优化开发的方法的适配,因此常依赖梯度代理,无法充分利用可用信息的零阶结构。本文提出一种用于非凸时变优化的随机两点直接搜索算法,并在恒定和递减探测率下推导迭代复杂度界。所得分析给出了关于问题时间变异性和可能的神谕误差的显式平稳性界。我们的复杂度界在时间不变设置下恢复了现有零阶方法的复杂度,同时将直接搜索方法扩展到静态设置之外。作为说明性应用,我们表明该方法自然适用于求解动力系统的最优(平衡选择)控制问题,在此设置下,分析给出了关于问题时间变异性的显式平稳性界,该变异性通过系统动态和外生扰动变化的影响来衡量。
英文摘要
Optimization under zeroth-order (i.e., bandit) feedback is central to many engineering problems where the analytic forms of objectives and/or constraints are unavailable. In modern applications, such as online control and online learning, optimization problems often evolve with time, requiring adaptive optimization methodologies. Yet, existing methods in this seting are largely confined to adaptations of methodologies developed for time-invariant or first-order optimization, and thus often rely on gradient surrogates that fail to fully exploit the zeroth-order structure of the available information. In this paper, we propose a randomized two-point direct-search algorithm for nonconvex time-varying optimization and derive iteration-complexity bounds under both constant and diminishing probing ratios. The resulting analysis yields explicit stationarity bounds in terms of the temporal variability of the problem and possible oracle errors. Our complexity bounds recover the complexity of existing zeroth-order methods in the time-invariant setting, while extending direct- search methods beyond static settings. As an illustrative application, we show that the methodology is naturally suited to solve optimal (equilibrium-selection) control problems for dynamical systems. In this setting, the analysis yields explicit stationarity bounds in terms of the temporal variability of the problem, measured through the effects of plant dynamics and exogenous disturbance variations.