arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10869cs.LGstat.ML

多类PAC学习的乐观速率

Optimistic Rates for Multiclass PAC Learning

Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究填补了多类PAC学习乐观速率的空白,确定了固定奥拉克风险下的最优超额风险界,扩展了相关理论至列表学习并改进了下界。

中文摘要 AI 辅助

当最优分类器已接近正确时,最坏情况多类界不会变小:缺失的是乐观速率,即一种波动随奥拉克风险本身缩放的保证。对于具有Natarajan维$d_N$和Daniely-Shalev-Shwartz维$d_{DS}$的类别,最优超额风险在两个端点处已知(可实现情况下为$d_{DS}/n$,不可知情况下为$\u221a{d_N/n}+d_{DS}/n$[HMZ24, CEH+26, Pab26]),中间区间仍是开放问题。我们填补了这一空白:在每个固定的奥拉克风险$L^\u2606$下,最优超额风险为$\u00d5(\u221a{L^\u2606 d_N/n}+d_{DS}/n)$,在字母表大小上一致,由一个既不知道$L^\u2606$也不知道置信水平的学习器达到。上界将[CEH+26]的覆盖-菜单-压缩架构(采用[Pab26]的可实现速率)与一个新的面向比较器的相对压缩定理相结合:一个经验上优于比较器$h$的大小为$k$的压缩规则,其总体风险至多为$L(h)+O(\u221a{L(h)\u0393}+\u0393)$,其中$\u0393=(k\ud835\udf52 n+\ud835\udf52(1/\u03b4))/n$,无需稳定性;这一结论传承了精确二元理论[MQZ26]的比较原理,同时摒弃了其无法推广至多类标签的布尔立方体几何结构。下界通过针对$L^\u2606$校准的成对Assouad方案,以及基于[BCD+22]中Natarajan维与DS维分离的伪立方体纤维论证,在每个固定$L^\u2606$下用一个类别和一个分布同时迫使两项存在。两个定理都扩展到列表学习:针对最优的$r$元假设组,相同的架构和相同的两个核心机制得出了形状相同的乐观速率和下界,证实了[Pab26]认为针对列表比较器必需的波动项,并从已知的可实现列表下界中移除了因子$r$。

英文摘要

Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales with the oracle risk itself. For a class of Natarajan dimension $d_N$ and Daniely-Shalev-Shwartz dimension $d_{DS}$, the optimal excess risk is known at the two endpoints ($d_{DS}/n$ realizable, $\sqrt{d_N/n}+d_{DS}/n$ agnostic [HMZ24, CEH+26, Pab26]) and open in between. We close the gap: at every fixed oracle risk $L^\star$, the optimal excess risk is $\widetildeΘ(\sqrt{L^\star d_N/n}+d_{DS}/n)$, uniformly in the alphabet size, attained by a learner that knows neither $L^\star$ nor the confidence level. The upper bound composes the cover-menu-compression architecture of [CEH+26], at the realizable rate of [Pab26], with a new comparator-facing relative compression theorem: a size-$k$ compression rule that empirically dominates a comparator $h$ has population risk at most $L(h)+O(\sqrt{L(h)Γ}+Γ)$ with $Γ=(k\log n+\log(1/δ))/n$, without stability; this transfers the comparison principle of the sharp binary theory [MQZ26] while discarding its Boolean-cube geometry, which does not lift to multiclass labels. The lower bound forces both terms using one class and one distribution at every fixed $L^\star$, by a pair-Assouad scheme calibrated to $L^\star$ and a fiber argument on the pseudo-cubes underlying the Natarajan-versus-DS separation of [BCD+22]. Both theorems extend to list learning: against the best $r$-tuple of hypotheses, the same architecture and the same two engines yield an optimistic rate and a lower bound of the same shape, forcing the fluctuation term that [Pab26] expected to be necessary against list comparators, and removing the factor $r$ from the known realizable list lower bound.

发表机构

  • University of New South Wales(新南威尔士大学)
  • University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑