arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Bi-ZOL:具有非光滑响应的双层零阶学习

Bi-ZOL: Bilevel Zeroth-Order Learning with Nonsmooth Responses

Zhisen Jiang, Saverio Bolognani

arXiv 2609.08021首次发表:更新:

发表机构

ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对下层约束双层优化中响应非光滑且模型信息不可用的问题,提出结构引导的零阶方法Bi-ZOL,分离链式法则结构估计超梯度,证明收敛性,实验显示优于普通零阶平滑。

AI 中文摘要

本文研究了响应预言机设置下的下层约束双层优化问题,其中下层模型信息不可用,且诱导的响应映射是局部Lipschitz的但可能非光滑。在此设置下,经典的响应雅可比矩阵和简化超梯度可能不存在。我们提出了双层零阶学习(Bi-ZOL),一种结构引导的零阶方法,用于寻找非光滑简化问题的驻点。Bi-ZOL不是估计完全平滑的简化超目标的梯度,而是分离双层链式法则结构:它保留在查询响应处的精确上层偏梯度,仅使用零阶采样来估计响应雅可比矩阵。这种构造产生了一个近似超梯度,它与Clarke链式法则次微分更直接地对齐。我们证明了Bi-ZOL方向具有部分平滑解释,量化了其逐点结构偏差,并证明了有限时间收敛到$(\u03b4,\u03b5)$-Bi-ZOL Frank-Wolfe驻点。对于分段$C^{1,1}$响应,在局部正则性下偏差为$O(\u03b4)$,对于活动单元邻域上的分段仿射响应,偏差消失。在基于激励的跟踪问题上的实验表明,在相当的响应预言机预算下,Bi-ZOL比普通零阶平滑实现了更小的驻点间隙和更低的超目标值。

英文摘要

This paper studies lower-level-constrained bilevel optimization in a response-oracle setting, where lower-level model information is unavailable and the induced response mapping is locally Lipschitz but potentially nonsmooth. In this setting, the classical response Jacobian and reduced hypergradient may fail to exist. We propose Bilevel Zeroth-Order Learning (Bi-ZOL), a structure-guided zeroth-order method for finding stationary points of the nonsmooth reduced problem. Instead of estimating the gradient of a fully smoothed reduced hyperobjective, Bi-ZOL separates the bilevel chain-rule structure: it keeps the exact upper-level partial gradients at the queried response and uses zeroth-order sampling only to estimate the response Jacobian. This construction yields an approximate hypergradient that is more directly aligned with the Clarke chain-rule subdifferential. We show that the Bi-ZOL direction admits a partial-smoothing interpretation, quantify its pointwise structural bias, and prove finite-time convergence to a $(δ,ε)$-Bi-ZOL Frank--Wolfe stationary point. The bias is $O(δ)$ for piecewise $C^{1,1}$ responses under local regularity and vanishes for piecewise affine responses on active-cell neighborhoods. Experiments on incentive-based tracking problems show that Bi-ZOL achieves smaller stationarity gaps and lower hyperobjective values than vanilla zeroth-order smoothing under comparable response-oracle budgets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑