arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10886cs.LGstat.ML

相对智能 II:可处理或半监督的实例最优学习

Relatively Smart II: Tractable or Semi-Supervised Instance-Optimal Learning

  • University of Southern California(南加州大学)
  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

Shaddin Dughmi, Alireza F. Pour

中文总结 AI 辅助

本研究证明ERM等适当一致学习器在无分布二分类中相对智能,并实现半监督相对智能学习,仅在无标签样本复杂度上有二次膨胀,但需超多项式预言机调用,牺牲可处理性。

中文摘要 AI 辅助

我们继续研究由Dughmi和Pour (2026)引入的相对智能学习,该学习要求监督学习者在每个边际上,与从无标签数据中可靠认证的每个分布固定误差保证竞争。他们证明了一包含图(OIG)学习器是相对智能的,但样本复杂度有二次方膨胀,且没有相对智能的学习器能做得更好,这留下了ERM或其他自然或可处理的学习器是否能实现类似保证的开放问题。他们还留下了膨胀是否能仅限于无标签数据的开放问题。我们的第一个结果表明,在无分布设置下,对于二分类,ERM——实际上任何适当的一致学习器——都是相对智能的。我们证明,用$m$个样本获得的小的可认证误差,意味着在大小为$O(m^2)$的随机样本上的均匀分布上也有类似的小误差,从而在该样本上产生一个大小至多为$2^{m+1}$的覆盖。这足以用$O(m^2)$个样本控制适当一致学习器的误差。然后我们证明,半监督相对智能学习在信息论上是可能的,仅在无标签样本复杂度上有二次方膨胀,而有标签样本复杂度无膨胀。该学习器使用OIG到一种留最多转导问题的自然推广,其中有限池中部分标签被揭示,其余标签被预测。最后,这种标签效率以简单性和可处理性为代价。如果假设类仅通过不可知ERM预言机访问,任何具有显著次二次方有标签样本膨胀的半监督相对智能学习器都需要超多项式次预言机调用。即使边际被明确给出,这也成立,因此也产生了分布固定学习的不可处理性结果,这可能具有独立意义。

英文摘要

We continue the study of relatively smart learning, introduced by Dughmi and Pour (2026), which asks a supervised learner to compete, marginal by marginal, with every distribution-fixed error guarantee soundly certifiable from unlabeled data. They showed that the One-Inclusion Graph (OIG) learner is relatively smart with a quadratic sample-complexity blowup, and that no relatively smart learner can do better, leaving open whether ERM or another natural or tractable learner achieves comparable guarantees. They also left open whether the blowup can be restricted to unlabeled data. Our firs results shows that ERM---and in fact any proper consistent learner---is relatively smart for binary classification in the distribution-free setting. We show that a small certifiable error with $m$ samples implies a similarly small error on the uniform distribution over a random sample of size $O(m^2)$, yielding a cover of size at most $2^{m+1}$ on that sample. This suffices to control the error of proper consistent learners with $O(m^2)$ samples. We then show that semi-supervised relatively smart learning is information-theoretically possible with a quadratic blowup only in unlabeled sample complexity and no blowup in labeled sample complexity. The learner uses a natural generalization of OIG to a leave-most-out transductive problem, where labels of part of a finite pool are revealed and the remaining labels are predicted. Finally, this label efficiency comes at a cost in simplicity and tractability. If the hypothesis class is accessed only through an agnostic ERM oracle, any semi-supervised relatively smart learner with substantially sub-quadratic labeled-sample blowup requires super-polynomially many oracle calls. This holds even when the marginal is given explicitly, and thus also yields an intractability result for distribution-fixed learning that may be of independent interest.

↑