arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通用可学习概念类的约束学习

Constrained Learning with Universally Learnable Concept Classes

Herlock SeyedAbolfazl Rahimi, Spyridon Pougkakiotis, Dionysis Kalogerias

arXiv 2608.08414首次发表:更新:

AI 中文总结

该研究在完全非凸设定下,建立无穷维假设类上对偶算法解的通用PACC学习性,引入闭包-实现间隙刻画可行性,证明最优值精确学习性,明确样本阈值为$1/\u03b5$的多项式,为大假设类上对偶学习提供规范框架。

AI 中文摘要

我们研究完全非凸设定下无穷维假设类上的约束统计学习问题,并建立对偶算法解的通用PACC学习性:即约束下的概率近似正确(Probably Approximately Correct on Constraints),该性质可同时保证最优性与约束满足。这强化了近PACC结果,后者的可行性残差是任何数据都无法消除的。最优性介于由Rademacher复杂度控制、倾向于小类别的泛化,以及依赖向量测度的Lyapunov凸性、需要可分解性(一种相反方向的要求)的强拉格朗日对偶之间。我们通过在通用RKHS $\u0397_K$(稠密于可分解包络)上建立总体问题,并在半径递增的范数球上学习来协调二者。这产生了Tikhonov复杂度 $\u1d4a8^\varepsilon_n$,即达到$\u03b5$-最优拉格朗日水平集的最小RKHS范数;我们证明其有限性,得到最优值的精确学习性,并在源条件下使样本阈值显式且为$1/\u03b5$的多项式。可行性更难:缺乏凸性时拉格朗日可能无法达到下确界,且对偶信息仅确定平均约束-风险向量,而非任何返回预测器的风险。我们引入闭包-实现间隙$\u03b5^\star_\infty$,它是$\u0397_K$通过对偶化检索可行解的指标,属于问题的性质而非建模选择。当$\u03b5^\star_\infty=0$时学习性精确,尤其在对偶可微性下;否则为近PACC,残差恰好为$\u03b5^\star_\infty$。最后,即使在无约束特例中也不存在无分布阈值,因此通用性是大假设类上对偶算法的规范框架。

英文摘要

We study constrained statistical learning over infinite-dimensional hypothesis classes in the fully nonconvex setting, and establish universal PACC learnability of the solutions of dual algorithms: Probably Approximately Correct on Constraints, guaranteeing optimality and constraint satisfaction at once. This strengthens near-PACC results, whose feasibility residual no amount of data can remove. Optimality is caught between generalization, governed by Rademacher complexity and favoring small classes, and strong Lagrangian duality, which rests on Lyapunov convexity for vector measures and needs decomposability, a demand pulling the other way. We reconcile the two by posing the population problem over a universal RKHS $\mathcal{H}_K$, dense in a decomposable envelope, and learning over norm balls of growing radius. This yields the Tikhonov complexity $\mathfrak{T}^{\varepsilon}_{n}$, the least RKHS norm reaching an $\varepsilon$-optimal Lagrangian level set; we prove it finite, obtain exact learnability of the optimal value, and make the sample threshold explicit and polynomial in $1/\varepsilon$ under a source condition. Feasibility is harder: absent convexity the Lagrangian may not attain its infimum, and dual information pins down only an averaged constraint-risk vector, not the risks of any returned predictor. We introduce the closure-realization gap $\varepsilon^\star_\infty$, an index of how well $\mathcal{H}_K$ retrieves feasible solutions from dualization; it is a property of the problem, not of a modeling choice. Learnability is exact when $\varepsilon^\star_\infty=0$, in particular under dual differentiability, and near-PACC with residual exactly $\varepsilon^\star_\infty$ otherwise. Finally, no distribution-free threshold exists already in the unconstrained specialization, so universality is the canonical frame for dual algorithms over large hypothesis classes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑