AI 中文总结
提出分解双层搜索(DBS),将变度量邻近梯度方法的缩放邻近步骤简化为二维残差系统,在SLOPE和组套索逻辑回归实验中,其效率和收敛性优于对比方法。
AI 中文摘要
变度量邻近方法可加速复合凸优化,但拟牛顿度量诱导的缩放邻近映射通常无闭式解。我们提出基于对角加秩一因子\boldsymbol{X}=\boldsymbol{D}+\boldsymbol{u}\boldsymbol{v}^\top的分解双层搜索(Decomposed Bilevel Search, DBS),其对角缩放满足弱割线方程。诱导的逆度量\boldsymbol{B}^{-1}=\boldsymbol{X}\boldsymbol{X}^\top可恢复零记忆DFP/BFGS型布罗伊登族成员,而因子形式将每个缩放邻近步骤简化为二维单调残差系统。每次残差评估需一次对角度量邻近映射,且经过认证的双层求解器可在\tilde{\boldsymbol{O}}((d+T_p)\boldsymbol{\nabla}^2(1/\boldsymbol{\nabla}))的计算量内达到目标精度\boldsymbol{\nabla},其中\boldsymbol{T}_p为该邻近映射的成本。在强凸性条件下,外层方法在精确和不精确内层求解时均线性收敛。在标量特例\boldsymbol{D}=\boldsymbol{\nabla}\boldsymbol{I}中,该预言机仅使用正则化项的普通邻近评估,无需广义雅可比或活动集信息。我们对有序加权\boldsymbol{\nabla}_1(SLOPE/OWL)和组套索逻辑回归开展实验,分别评估内层预言机和完整外层方法。该预言机在高维SLOPE缩放邻近子问题(条件数高达\boldsymbol{4\times10^6})上仅用数百次普通邻近评估即可求解;在SLOPE外层基准测试中,带热启动的预言机可控制缩放邻近开销,且DBS达到严格目标所需的梯度评估次数远少于 Lipschitz 归一化FISTA,在\texttt{real-sim}数据集上表现出明显的目标时间优势;在组套索逻辑回归上,DBS在合成相关实例中表现可靠,在真实数据分组复制实例中速度最快。
英文摘要
Variable-metric proximal methods accelerate composite convex optimization, but the scaled proximal map induced by a quasi-Newton metric rarely has a closed form. We develop \emph{Decomposed Bilevel Search} (DBS), based on a diagonal-plus-rank-one factor \(X=D+uv^\top\) whose diagonal scaling satisfies a weak secant equation. The induced inverse metric \(B^{-1}=XX^\top\) recovers zero-memory DFP/BFGS-type Broyden members, while the factor form reduces each scaled proximal step to a two-dimensional monotone residual system. Each residual evaluation requires one diagonal-metric proximal map, and a certified bilevel solve reaches target accuracy \(ε\) in \(\mathcal O((d+T_p)\log^2(1/ε))\) work, where \(T_p\) is the cost of that proximal map. Under strong convexity, the outer method converges linearly with exact and inexact inner solves. In the scalar specialization \(D=αI\), the oracle uses only ordinary proximal evaluations of the regularizer and no generalized Jacobian or active-set information. Experiments on ordered-weighted \(\ell_1\) (SLOPE/OWL) and group-lasso logistic regression evaluate both the inner oracle and the full outer method. The oracle solves high-dimensional SLOPE scaled proximal subproblems of condition number up to \(4\times10^6\) using a few hundred ordinary proximal evaluations. In SLOPE outer benchmarks, the warm-started oracle keeps the scaled-proximal overhead controlled and DBS reaches stringent targets with far fewer gradient evaluations than Lipschitz-normalized FISTA; on \texttt{real-sim} this becomes a clear target-time advantage. On group-lasso logistic regression, DBS is reliable on synthetic correlated instances and fastest on a real-data grouped-copy instance.
Comments37 pages, 1 figure