AI 中文总结
本研究修正了 Bandit 多类 PAC 学习的下界,引入锚定维数 aBDS,移除标签数 K 的上界依赖,并发现置信度直和现象,表明仅凭维数无法精确刻画样本复杂度。
AI 中文摘要
我们研究带 bandit 反馈的可实现多类 PAC 学习:学习器观察一个 i.i.d. 实例,预测 $K$ 个标签中的一个,并且仅得知预测是否正确。Hanneke、Meng、Moran 和 Shaeiri(arXiv:2605.25678)通过 bandit DS 维数 $\mathrm{BDS}$ 在对数因子范围内刻画了最优样本复杂度,并询问是否每个类别都承认样本复杂度 $O((\mathrm{BDS}+\log(1/\delta))/\epsilon)$。首先,我们表明已发表的下界 $\Omega((\mathrm{BDS}+\log(1/\delta))/\epsilon)$ 按原样陈述是错误的:我们展示了具有 $\mathrm{BDS}=K-1$ 的显式类别,其样本复杂度指数级更小,并在其证明中定位了两个独立的缺口。我们围绕一个新的锚定维数 $\mathrm{aBDS}\le\mathrm{BDS}$ 修复了下界理论,证明了一个无常数的三部分下界。在上界方面,我们完全移除了环境标签数量 $K$,证明了 $O((B\log^3 B+B\log(1/\delta))/\epsilon)$,其中 $B=\mathrm{BDS}$,并通过一个新的纤维化引理得到了一个常数置信度界;对于两个自然族,我们确定了样本复杂度直至常数因子。最后,我们在其均匀常数解读下否定了开放问题,并表明失败是内在的:对于一个显式的仿射多路复用器类别,我们建立了完整的置信度轮廓 $\Theta((n\min{n,\log(1/\delta)}+\log(1/\delta))/\epsilon)$,这是一个置信度直和机制,其中乘性的 $\log(1/\delta)$ 成本在信息论上是必要的,随后是秩饱和相变。两个具有相同维数轮廓的类别可以具有多项式不同的样本复杂度,因此仅凭这些维数的任何刻画都不能精确到多对数因子。我们还表明这些结果与通过 ListCascade 桥接的加性置信度列表-PAC 保证一致。
英文摘要
We study realizable multiclass PAC learning with bandit feedback: the learner observes an i.i.d. instance, predicts one of $K$ labels, and learns only whether the prediction was correct. Hanneke, Meng, Moran, and Shaeiri (arXiv:2605.25678) characterized the optimal sample complexity via the bandit DS dimension $\mathrm{BDS}$ up to logarithmic factors, and asked whether every class admits sample complexity $O((\mathrm{BDS}+\log(1/δ))/ε)$. First, we show that the published lower bound $Ω((\mathrm{BDS}+\log(1/δ))/ε)$ is incorrect as stated: we exhibit explicit classes with $\mathrm{BDS}=K-1$ whose sample complexity is exponentially smaller, and locate two independent gaps in its proof. We repair the lower-bound theory around a new anchored dimension $\mathrm{aBDS}\le\mathrm{BDS}$, proving a constant-free three-part lower bound. On the upper-bound side we remove the ambient label count $K$ entirely, proving $O((B\log^3 B+B\log(1/δ))/ε)$ for $B=\mathrm{BDS}$, plus a constant-confidence bound via a new fiberization lemma; for two natural families we determine the sample complexity up to constant factors. Finally, we answer the open question in the negative under its uniform-constant reading, and show the failure is intrinsic: for an explicit affine multiplexer class we establish the full confidence profile $Θ((n\min{n,\log(1/δ)}+\log(1/δ))/ε)$, a confidence direct-sum regime where a multiplicative $\log(1/δ)$ cost is information-theoretically necessary, followed by a rank-saturation phase transition. Two classes with identical dimension profiles can have polynomially different sample complexities, so no characterization by these dimensions alone is accurate to polylogarithmic factors. We also show these results are consistent with additive-confidence list-PAC guarantees via the ListCascade bridge.
Comments23 pages