发表机构
Department of Statistics, University of Oxford; Department of Computer Science, ETH Zurich(牛津大学统计学系; 苏黎世联邦理工学院计算机系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对回归问题,证明有限假设类下$Q$-聚合估计量可同时实现极小极大与通用指数速率,可数无限假设类下二者存在固有权衡,还给出平方损失学习通用速率的结构性结果。
AI 中文摘要
我们研究带界响应下的回归问题,以 excess mean squared error( excess 均方误差)为评价指标。当比较类有限时,该设定被称为模型选择聚合,要达到 minimax excess risk(极小极大 excess 风险)需使用不当学习算法。与之相反,在通用学习框架中无需不当学习,简单的经验风险最小化(ERM)即可达到最优指数学习速率。因此,两种框架指向不同的最优算法原则,这引出了两全其美的保证问题:极小极大速率与通用指数速率能否由同一算法实现?对于有限假设类,我们给出肯定回答,证明$Q$-聚合估计量(已知其可达到极小极大最优尾概率)可实现指数通用速率;而大量其他估计量和算法原则(ERM、序列平均、剪枝、星型估计)无法同时达到这两种速率。对于可数无限假设类,我们给出否定回答,证明在达到指数通用速率与极小极大均匀速率之间存在固有权衡,且该权衡可通过结合两种框架的最优算法,利用$Q$-聚合精确刻画。此外,我们还证明了平方损失学习中通用速率的若干额外结构性结果。
英文摘要
We study regression under bounded responses in terms of excess mean squared error. When the comparator class is finite, this setting is known as model selection aggregation, and achieving minimax excess risk requires improper learning algorithms. Contrary to this, in the universal learning framework no improperness is needed, as simple empirical risk minimization achieves the best-possible exponential learning rate. Hence, the two frameworks suggest different optimal algorithmic principles. This poses the question of best-of-both-worlds guarantees: Are minimax and universal exponential rates achievable by the same algorithm? For finite hypothesis classes, we answer this question in the affirmative by showing that the $Q$-aggregation estimator - which is known to achieve minimax optimal tails - achieves exponential universal rates. A wide range of other estimators and algorithmic principles (ERM, sequential averaging, pruning, and star estimation) do not achieve both. For countably infinite hypothesis classes, we answer the question in the negative by showing that there is an inherent trade-off between achieving exponential universal and minimax uniform rates. This trade-off is exactly traced by combining optimal algorithms from each world using $Q$-aggregation. Besides these results, we prove several additional structural results about universal rates in learning with squared loss.