arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08416cs.LG

Tsybakov噪声下的最优学习

Optimal Learning Under Tsybakov Noise

Steve Hanneke, Hongao Wang, Mingyue Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究解决了Tsybakov噪声下PAC学习的开放问题,通过改进上界匹配已知最优下界确立最优误差保证,算法自适应划分实例空间并输出满足区域误差约束的假设。

中文摘要 AI 辅助

可能近似正确(PAC)学习[Val84]是一种被广泛研究的基础学习模型。在该模型中,$\boldsymbol{\textit{H}} \boldsymbol{\boldsymbol{\text{⊆}}} \boldsymbol{\textit{\text{X}}}$是概念类,$\boldsymbol{h^* \boldsymbol{\boldsymbol{\text{∈}}} \boldsymbol{\textit{H}}}$是待学习的目标概念。通过获取来自$\boldsymbol{\text{X}×{0,1}}$上分布$\boldsymbol{\text{D}}$的独立同分布标记样本(其中$\boldsymbol{h^*}$是$\boldsymbol{\text{H}}$中最优概念),目标是设计一种学习算法,使其输出的假设具有低误差,且以高概率与$\boldsymbol{h^*}$的误差相当。该模型最初在可实现设定下研究,该设定假设$\boldsymbol{h^*}$无误差。一种自然的放松是允许标签噪声,即真实标签以概率$\boldsymbol{\text{η∈(0,1/2)}}$翻转。现实中,某些标签可能噪声极大,尤其是决策边界附近的点,因此允许极少出现的极噪声点是自然的,这由[MT99]和[Tsy04]引入的噪声模型量化,即现在所知的Tsybakov噪声。对于一般概念类的学习,[MN06]给出了Tsybakov噪声下误差保证的一般上界和下界,但它们相差一个对数因子,解决该差距是过去二十年的著名开放问题。本研究通过将上界改进至匹配已知最优下界,解决了该开放问题,从而确立了Tsybakov噪声下学习的最优误差保证。所提学习算法通过将实例空间自适应划分为大致对应不同噪声水平的区域,并为每个区域返回满足特定误差约束的概念类中的假设来运行,该技术与非可实现学习的若干近期进展(如[HLZ24]和[Han25])具有概念基础的共通性。

英文摘要

Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated. In this model, $\mathcal{H} \subseteq \{0,1\}^{\mathcal{X}}$ is a concept class, and $h^*\in\mathcal{H}$ is the target concept to be learned. Having access to i.i.d. labeled examples from a distribution $\mathcal{D}$ over $\mathcal{X}\times\{0,1\}$, which admits $h^*$ as the best concept in $\mathcal{H}$, the goal is to design a learning algorithm that outputs a hypothesis having low error competitive to $h^{*}$ with high probability. This model was initially studied under the realizable setting, which assumes that $h^*$ has no error. A natural relaxation is to allow label noise, that is, the true label can be flipped with probability $η\in(0,1/2)$. In reality, certain labels might be extremely noisy, especially for those points near the decision boundary. Hence, it is natural to allow very noisy points, though only rarely. This is quantified by a noise model introduced by [MT99] and [Tsy04], now known as Tsybakov noise. For learning general concept classes, [MN06] gave the general upper and lower bounds for error guarantees under Tsybakov noise. However, their upper and lower bounds differ by a logarithmic factor. Resolving this gap has remained a well-known open question for the past twenty years. In this work, we resolve this open question by improving the upper bound to match the best known lower bound, thus establishing the optimal error guarantee for learning under Tsybakov noise. Our learning algorithm operates by adaptively partitioning the instance space into regions, roughly corresponding to different noise levels, and returning a hypothesis in the concept class satisfying a specific error constraint for each region. Our technique shares a conceptual foundation with several recent advances in non-realizable learning, such as [HLZ24] and [Han25].

发表机构

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑