任意维数严格凸二次型循环最速下降法的双重指数收敛性
Doubly exponential convergence of the cyclic steepest descent method for strictly convex quadratics in arbitrary dimensions
浏览论文内容
中文总结 AI 辅助
本文证明任意维数下,当不同特征值数小于两倍循环长度时,循环最速下降法具有双重指数收敛,首次给出严格证明并确认阈值。
中文摘要 AI 辅助
我们研究严格凸二次型最小化的循环最速下降法(CSD)。CSD 在每个循环开始时计算精确的最速下降步长,并在连续 j 次迭代中重复使用该步长。迄今为止,针对真实算法的唯一严格结果仅限于二维情形,在该情形下已知循环起始梯度序列双重指数收敛。对于通过丢弃有界对数项得到的简化(“简单”)模型,Dai 和 Fletcher 预测:当不同特征值的个数 n 小于循环长度 m 的两倍时,CSD 是超线性收敛的,否则为线性收敛。设 A 为对称正定矩阵,具有 q 个不同特征值,并假设 2j > q。我们证明了在任意维数下真实算法在该阈值下的超线性侧,实际上获得了更快的双重指数衰减。除了一组勒贝格零测度的初始点外,对于低于显式正阈值 kappa_0 的每个 kappa,循环起始梯度满足 ||g_k|| <= exp(-C exp(kappa k)),并且完整梯度和迭代误差在 m 次迭代后满足具有指数 kappa/j 的类似界。这是对任意维数(包括重特征值)下真实 CSD 的首次严格证明,并证实了阈值 2j > q(简单模型预测 n < 2m);双重指数速率比简单模型预测的超线性速率更快。证明结合了任意参考比率坐标、逆像体积收缩、薄带估计和逐分量 Borel-Cantelli 论证。
英文摘要
We study the cyclic steepest descent method (CSD) for strictly convex quadratic minimization. CSD repeats, for j consecutive iterations, the exact steepest-descent step size computed at the beginning of each cycle. The only rigorous result for the real algorithm has so far been restricted to two dimensions, where the cycle-starting gradient sequence is known to converge doubly exponentially. For the simplified ("simple") model obtained by discarding bounded logarithmic terms, Dai and Fletcher predicted that CSD is superlinear whenever the number n of distinct eigenvalues is below twice the cycle length m, and linear otherwise. Let A be symmetric positive definite with q distinct eigenvalues, and suppose 2j > q. We prove the superlinear side of this threshold for the real algorithm in arbitrary dimensions, and in fact obtain a faster, doubly exponential, decay. Except for a Lebesgue-null set of initial points, for every kappa below an explicit positive threshold kappa_0, the cycle-starting gradient satisfies ||g_k|| <= exp(-C exp(kappa k)), and the full gradient and iterate errors satisfy analogous bounds with exponent kappa/j after m iterations. This is the first rigorous proof for the real CSD in arbitrary dimensions, including repeated eigenvalues, and confirms the threshold 2j > q (the simple-model prediction n < 2m); the doubly exponential rate is faster than the superlinear one predicted by the simple model. The proof combines arbitrary-reference ratio coordinates, inverse-image volume contraction, thin-band estimates, and a per-component Borel-Cantelli argument.
发表机构
- Nankai University(南开大学)
机构由 AI 辅助整理,请以论文原文为准。