发表机构
CERMICS, CNRS, ENPC, Institut Polytechnique de Paris; Institut Camille Jordan, École Centrale Lyon, CNRS UMR 5208; Institut Universitaire de France (IUF)(法国国家科学研究中心,国立桥路学校,巴黎理工学院; 里昂中央理工学院,法国国家科学研究中心; 法国高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于 Fenchel-Young 对偶间隙的可计算误差界与提前停止准则,用于正则化逆问题,并应用于广义 Beurling-Lasso 及深度学习优化器认证。
AI 中文摘要
我们研究正则化逆问题的可计算误差界和认证提前停止,其中数据保真项与正则化项进行权衡。分析依赖于一个精确的对偶间隙恒等式,该恒等式将 $F(\Phi\mu)+\lambda R(\mu)$ 的总间隙分解为数据保真 Fenchel-Young 损失和正则化器 Fenchel-Young 损失,$\Delta(\mu,h)=L_F(\Phi\mu\parallel h)+\lambda L_R(\mu\parallel\eta),\qquad \eta=-\Phi^\star h/\lambda$,该恒等式对任何原始点 $\mu$ 和任何对偶点 $h$ 均成立。数据保真项 $F$ 是严格凸的,因此在 $F^\star$ 可微的任何地方,损失 $L_F(\Phi\mu\parallel h)$ 是 $F$ 在预测 $\Phi\mu$ 和 $\nabla F^\star(h)$ 之间的 Bregman 散度,并且当且仅当满足镜像对齐 $h=\nabla F(\Phi\mu)$ 时它恰好为零。在对偶可行点 $\tilde h$ 处评估时,间隙 $\Delta(\mu,\tilde h)$ 是可计算的且无预言机,这意味着它不使用解的任何知识,并且它界定了 $\mu$ 的次优性。在源条件下,相同的 Fenchel-Young 损失给出了估计误差和预测误差的先验界。它们的尺度是不可约的模型和噪声误差 $L_F(\Phi\mu^\star\parallel h^\star)$,当证书处满足镜像对齐时该误差恰好为零。这给出了一个提前停止规则:运行算法直到正则化器 Fenchel-Young 损失低于容差 $\epsilon$。Brøndsted-Rockafellar 定理的一个构造性版本随后将当前对转换为精确的对偶可行对,并且该代理提升为精确的原始证书。我们通过在保真度几何中进行近端步骤来构建代理,使用 Bregman 核 $F^\star$ 并由预测 $\Phi\mu$ 倾斜:当 $F$ 是平方误差时,它恢复了 Carlier 的欧几里得步骤,并且它将对偶间隙减少了正则化器 Fenchel-Young 损失,直到一个在二次情形下消失的二阶余项。我们的运行示例是广义 Beurling-Lasso (GBL),其中 $R$ 是带符号测度上的全变差范数。它包含经典的 Beurling-Lasso(通过平方误差获得),并且还涵盖了鲁棒、逻辑、熵和逆最优传输损失。相同的对偶间隙认证了深度学习优化器(如 Lion-K 和 Muon)以其近端形式作为正则化程序的求解器。同一作者的一篇配套论文基于这些误差界,在非退化源条件下建立了 GBL 的精确支持恢复。
英文摘要
We study computable error bounds and certified early stopping for regularized inverse problems, where a data-fidelity term is traded against a regularizer. The analysis relies on an exact duality-gap identity that splits the total gap of $F(Φμ)+λR(μ)$ into a data-fidelity Fenchel--Young loss and a regularizer Fenchel--Young loss, $ Δ(μ,h)=L_F(Φμ\parallel h)+λL_R(μ\parallelη),\qquad η=-Φ^\star h/λ, $ valid for any primal point $μ$ and any dual point $h$. The data-fidelity term $F$ is strictly convex, so wherever $F^\star$ is differentiable the loss $L_F(Φμ\parallel h)$ is the Bregman divergence of~$F$ between the prediction $Φμ$ and $\nabla F^\star(h)$, and it vanishes exactly at \emph{Mirror Alignment} $h=\nabla F(Φμ)$. Evaluated at a dual-feasible point $\tilde h$, the gap~$Δ(μ,\tilde h)$ is computable and \emph{oracle-free}, meaning that it uses no knowledge of the solution, and it bounds the suboptimality of $μ$. Under the source condition, the same Fenchel--Young losses give \emph{a priori} bounds on the estimation and prediction errors. Their scale is the irreducible model and noise error $L_F(Φμ^\star\parallel h^\star)$, which vanishes exactly when Mirror Alignment holds at the certificate. This gives an early-stopping rule: run the algorithm until the regularizer Fenchel--Young loss falls below a tolerance $ε$. A constructive version of the Brøndsted--Rockafellar theorem then turns the current pair into an exact \emph{dual-feasible} one, and this proxy lifts to an exact primal certificate. We build the proxy by a proximal step in the geometry of the fidelity, with Bregman kernel $F^\star$ and tilted by the prediction $Φμ$: it recovers the Euclidean step of Carlier when $F$ is the squared error, and it reduces the duality gap by the regularizer Fenchel--Young loss, up to a second-order remainder that vanishes in the quadratic case. Our running example is the Generalized Beurling--Lasso (GBL), where $R$ is the total-variation norm on signed measures. It contains the classical Beurling--Lasso, obtained with the squared error, and also covers robust, logistic, entropic and inverse-optimal-transport losses. The same duality gap certifies deep-learning optimizers such as Lion-K and Muon, in their proximal form, as solvers of the regularized program. A companion paper by the same authors builds on these error bounds to establish exact support recovery for the GBL under a non-degenerate source condition.