发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对公共随机损失下的随机二阶平稳性,我们建立了匹配的极小极大容差界,上界移除混合项,下界通过光滑随机损失实现,并推广至Hölder Hessian情形。
AI 中文摘要
我们建立了随机二阶平稳性的紧多项式容差界,其中每次新的预言机响应是公共随机标量损失的导数。对于具有Lipschitz梯度和Hessian的总体目标$F$,目标是$\n|\nabla F(x)\\|\n≤ ε$和$λ_{\min}(\nabla^2 F(x))≥ -γ$,其中容差$ε,γ>0$独立。在有界梯度方差和几乎必然有界Hessian误差下,新的梯度或Hessian-向量积调用的极小极大次数为$\widetildeΘ\left(ε^{-3}+γ^{-5}\right)$。该刻画固定了正间隙、光滑性和噪声参数,抑制了对数因子,并允许维度在显式多项式包络内增长。上界从先前的new-HVP保证中移除了混合项$ε^{-2}γ^{-2}$。直接随机线Hessian估计和二元梯度跟踪器将梯度漂移与随机符号曲率运动分离。下界通过全局定义的光滑随机损失实现端点成本:光滑划分将标量噪声局部化而无链长惩罚,而精确抵消限制了整个响应的信息。因此,即使对于具有有界值方差的联合值、梯度和全Hessian响应,相同的容差指数也成立。对于具有Hölder指数$ν\in(0,1]$的总体Hessian,在相应的维度包络下,新的梯度/HVP复杂度变为$\widetildeΘ\left(ε^{-3}+γ^{-(3+2/ν)}\right)$。
英文摘要
We establish tight polynomial tolerance bounds for stochastic second-order stationarity when each fresh oracle response is a derivative of one common random scalar loss. For a population objective $F$ with Lipschitz gradient and Hessian, the target is $\|\nabla F(x)\|\le ε$ and $λ_{\min}(\nabla^2 F(x))\ge -γ$, with independent tolerances $ε,γ>0$. Under bounded gradient variance and almost-surely bounded Hessian error, the minimax number of fresh gradient or Hessian-vector-product calls is $\widetildeΘ\!\left(ε^{-3}+γ^{-5}\right)$. The characterization fixes positive gap, smoothness, and noise parameters, suppresses logarithmic factors, and allows dimension to grow within an explicit polynomial envelope. The upper bound removes the mixed term $ε^{-2}γ^{-2}$ from the earlier fresh-HVP guarantee. Direct random-line Hessian estimates and a dyadic gradient tracker separate gradient drift from randomly signed curvature motion. The lower bound realizes the endpoint costs through globally defined smooth random losses: a smooth partition localizes scalar noise without a chain-length penalty, while exact cancellation limits the information in the entire response. Consequently, the same tolerance exponents hold even for joint value, gradient, and full-Hessian responses with bounded value variance. For a population Hessian with Hölder exponent $ν\in(0,1]$, fresh gradient/HVP complexity becomes $\widetildeΘ\!\left(ε^{-3}+γ^{-(3+2/ν)}\right)$ under the corresponding dimension envelope.