关于带成员查询的主动学习样本复杂度的研究
On the Sample Complexity of Active Learning with Membership Queries
- University of Arizona(亚利桑那大学)
- University of Southern California(南加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究主动学习中合成查询的能力,发现其能显著改变学习难度,使某些类别从多项式误差衰减变为指数级可学习,并提出了充分条件与推测性视角。
AI中文摘要:
这项工作重新审视了主动学习中的一个基本问题:合成任意查询的能力有多强大?与基于池的主动学习(学习者仅从给定的未标记池中选择查询)相比,我们发现查询能力的这一看似微小的改变可能会显著改变统计学习的难度。特别是,某些在基于池的设置中固有地难以学习的假设类别,仅能实现样本数量上的多项式误差衰减,而一旦允许合成查询,它们就变得可以指数级地学习。这一显著差距表明,成员查询合成引发了一种根本不同的学习模式,这种模式未被现有的主动学习理论充分捕捉,需要新的分析工具来刻画其复杂性。受此现象启发,我们发展了几种充分条件,提供了有趣的例子,并提出了一个推测性的视角,以理解哪些假设类别可以通过合成查询实现高效学习。
英文摘要:
This work revisits a fundamental question in active learning: how powerful is the ability to synthesize arbitrary queries? Compared to pool-based active learning, where the learner only selects queries from a given unlabeled pool, we find that this seemingly mild change in query ability may dramatically alter the difficulty of statistical learning. In particular, some hypothesis classes that are inherently slow to learn in the pool-based setting, achieving only polynomial error decay in the number of samples, become exponentially learnable once synthesized queries are allowed. This striking gap suggests that membership query synthesis induces a fundamentally different mode of learning, one that is not adequately captured by existing active learning theory and calls for new analytical tools to characterize its complexity. Motivated by this phenomenon, we develop several sufficient conditions, present intriguing examples, and propose a conjectural perspective toward understanding which hypothesis classes admit efficient learning through synthesized queries.