发表机构
Ariel University(阿里尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对受Cressie–Read散度约束的分布鲁棒PAC学习,推导了可实现与不可知情况下的样本复杂度界,分析了鲁棒性对ε依赖关系的影响,扩展了χ²散度的相关结果并填补了界的空白。
AI 中文摘要
我们针对0-1损失研究分布鲁棒PAC学习,其中数据分布的对抗性扰动受阶数k>1、半径ρ≥0的Cressie–Read散度约束。对于VC维为d的假设类,我们分别建立了可实现和不可知情况下的样本复杂度界,前者的界在常数因子内是紧的,后者在对数因子内是紧的;普通经验风险最小化方法在对数因子内可达到这两种速率。对于目标精度ε∈(0,1)和置信度δ∈(0,1),它们的阶数分别为max{1/ε, ρ^(1/(k-1))/ε^(k_⋆)}·(d+logδ⁻¹)和max{1/ε², ρ^(1/(k-1))/ε^(k_⋆∨2)}·(d+logδ⁻¹),其中k_⋆=k/(k-1)。对于每个固定的ρ>0,当ε趋近于0时,鲁棒性将可实现情况下的ε依赖关系从ε⁻¹变为ε⁻k_⋆。在不可知情况下,当1<k<2时,鲁棒性将ε依赖关系从ε⁻²变为ε⁻k_⋆;而当k≥2时,指数保持经典的2,同时存在非平凡的ρ依赖关系。基于鲁棒0-1风险到普通分类误差的已知标量归约,我们的分析揭示了分类误差的统计估计与鲁棒性对其的放大之间的尺度敏感相互作用,清晰解释了不可知情况下速率的转变。我们将先前研究的χ²散度情况扩展到所有阶数k>1的Cressie–Read散度,填补了其上界和下界的空白,并且当ρ趋近于0时,我们的结果恢复了标准的PAC学习速率,这与先前无法在该极限下正确插值的边界不同。
英文摘要
We study distributionally robust PAC learning for the $0$--$1$-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--Read divergence of order $k>1$ and radius $ρ\geq 0$. For hypothesis classes with VC dimension $d$, we establish realizable and agnostic sample-complexity bounds tight up to constant and logarithmic factors, respectively; ordinary empirical risk minimization attains both rates up to logarithmic factors. For target accuracy $\varepsilon\in(0,1)$ and confidence $δ\in(0,1)$, their respective orders are \[ \max\!\left\{\frac{1}{\varepsilon}, \frac{ρ^{\frac 1{k-1}}}{\varepsilon^{k_\star}} \right\}\cdot(d+\log δ^{-1}) \qquad\text{and}\qquad \max\!\left\{\frac{1}{\varepsilon^2}, \frac{ρ^{\frac1{k-1}}}{\varepsilon^{k_\star\vee 2}} \right\}\cdot(d+\log δ^{-1}), \] where $k_\star={k}/{(k-1)}$. For every fixed $ρ>0$, robustness changes the realizable $\varepsilon$-dependence from $\varepsilon^{-1}$ to $\varepsilon^{-k_\star}$ as $\varepsilon\downarrow0$. In the agnostic case, for $1<k<2$, robustness changes the $\varepsilon$-dependence from $\varepsilon^{-2}$ to $\varepsilon^{-k_\star}$, whereas for $k\geq2$ the exponent remains the classical $2$, with nontrivial $ρ$-dependence. Building on the known scalar reduction of robust $0$--$1$ risk to ordinary classification error, our analysis reveals a scale-sensitive interaction between the statistical estimation of classification error and its amplification by robustness, sharply explaining the transition in the agnostic rate. We extend the previously studied $χ^2$-divergence case to every Cressie--Read order $k>1$, close its upper--lower gaps, and recover standard PAC learning rates as $ρ\to0$, unlike previous bounds that fail to interpolate correctly in this limit.