多分类学习的算法原理难以捉摸:正则化与恰当学习的局限性
Algorithmic Principles For Multiclass Learning Are Hard To Come By: Limits of Regularization and Proper Learning
浏览论文内容
中文总结 AI 辅助
该研究解决统计学习理论中三个开放问题,证明多分类学习无法简化为恰当学习,恰当学习存在亚线性误差必要条件,且正则化不是通用学习器,同时给出SRM可学习性的充分条件。
中文摘要 AI 辅助
统计学习理论中两个最基础的问题是:哪些预测问题是可学习的,以及应当如何学习它们。对于前者,简洁的答案通常以组合维数的形式呈现;然而后者却更难把握:所有已知的通用多分类学习器都依赖于指数级大一包含结构的复杂定向,而恰当学习(proper learning)和正则化这类熟悉的算法原理仍未被充分理解。受先前工作启发,我们提出问题:学习是否可简化为恰当学习(可能在更大的假设类上进行),以及恰当或非恰当多分类学习最终是否可由合适的正则项捕获?我们的主要结果对这两个问题给出否定回答,解决了先前工作中的三个开放问题。第一,我们展示了一个可学习的多分类问题,它无法嵌入任何可恰当学习的类中,即学习无法通过扩大假设类简化为恰当学习。第二,我们证明恰当学习可能需要训练误差,并精确描述了这一现象:每个可恰当学习的类都存在一个恰当学习器,在大小为m的样本上做出o(m)个错误,但对于某些可恰当学习的问题,每个规定的亚线性尺度a_m=o(m)都是必要的。第三,正则化不是通用学习器:我们展示了一个可恰当学习的类,它无法被任何结构风险最小化(Structural Risk Minimization, SRM)学习器学习,还有一个可学习的类,它无法被任何局部正则化器学习。我们补充这些不可能性结果的是一个正向理论,它给出了SRM可学习性的两个充分条件,并通过揭示偏好的可积性来刻画SRM可表示性。
英文摘要
Two of the most fundamental questions in statistical learning theory are the following: which prediction problems are learnable, and how should they be learned? For the former, elegant answers often take the form of combinatorial dimensions. The latter question, however, has proved considerably more elusive: all known general-purpose multiclass learners rely on intricate orientations of exponentially large one-inclusion structures, and familiar algorithmic principles such as proper learning and regularization remain poorly understood. Motivated by prior work, we ask whether learning reduces to proper learning---possibly over a larger hypothesis class---and whether proper or improper multiclass learning can ultimately be captured by suitable regularizers. Our primary results answer both questions negatively, resolving three open problems from prior work. First, we exhibit a learnable multiclass problem that cannot be embedded in any properly learnable class, meaning learning cannot be reduced to proper learning by enlarging the hypothesis class. Second, we demonstrate that proper learning can require training error and characterize this phenomenon precisely: every properly learnable class admits a proper learner making $o(m)$ errors on samples of size $m$, but every prescribed sublinear scale $a_m=o(m)$ is necessary for some properly learnable problem. Third, regularization is not a general learner: we exhibit a properly learnable class that cannot be learned by any Structural Risk Minimization (SRM) learner, and a learnable class that cannot be learned by any local regularizer. We complement these impossibility results with a positive theory that gives two sufficient conditions for SRM learnability and characterizes SRM representability through integrability of revealed preferences.
发表机构
- University of Chicago(芝加哥大学)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。