AI 中文总结
本研究提出防御提升器算法,可同时满足在线梯度提升与在线弱到强提升的两种保证,效率高且在合成和真实数据流上预测性能强劲、运行速度快。
AI 中文摘要
我们研究由自适应对手选择的二元结果的在线概率预测问题。给定一个针对弱假设类$H$的在线学习算法,我们希望高效获得现有在线提升技术分别提供的两种不可比的保证。在线梯度提升在每个序列上与$H$张成空间诱导的最佳预测器竞争布雷分数,但当该张成空间不包含准确预测器时则无法提供任何保证。在线弱到强提升在弱学习条件下将分类误差降至零,但当该条件不满足时则几乎无法提供保证。我们提出一种简单的防御性预测算法——防御提升器(Defensive Booster),可同时获得两种保证。在每个自适应序列上,其布雷分数以与在线梯度提升相同的速率与$H$张成空间诱导的最佳预测器竞争;同时,每当实现的转录满足平滑弱学习条件时,其布雷分数和随机分类误差满足与在线分类提升相同的速率保证。这通过实现提升的“对偶视角”实现:当算法的随机分类误差持续较高时,其错误权重形成平滑重加权,使得每个弱假设具有低边,从而产生事后核心证明,表明弱学习条件不满足。我们还开发了一种强自适应变体,在每个时间区间上满足两种保证。防御提升器效率极高:仅访问一个弱学习器,而我们对比的现有在线提升方法则维护大型弱学习器集成。在合成和真实数据流上的实验表明,其预测性能强劲(有时显著优于所有现有基线),同时运行速度快几个数量级。
英文摘要
We study online probabilistic forecasting of binary outcomes chosen by an adaptive adversary. Given an online learning algorithm for a weak hypothesis class $H$, we would like to efficiently obtain two incomparable guarantees that existing online boosting techniques provide separately. Online gradient boosting competes in Brier score with the best predictor induced by the span of $H$ on every sequence, but promises nothing when the span does not contain an accurate predictor. Online weak-to-strong boosting drives classification error to zero under a weak-learning condition, but promises little when that condition fails. We give a simple defensive forecasting algorithm, the Defensive Booster, that obtains both guarantees. On every adaptive sequence, its Brier score is competitive with the best prediction induced by the span of $H$ at the same rate as online gradient boosting; simultaneously, whenever the realized transcript satisfies the smooth weak-learning condition, its Brier score and randomized classification error satisfy the same rate guarantee as online classification boosting. This is achieved by operationalizing the "dual view" of boosting: When the algorithm's randomized classification error is persistently high, its mistake weights form a smooth reweighting on which every weak hypothesis has low edge, yielding an ex-post hard-core certificate that the weak-learning condition fails. We also develop a strongly adaptive variant, which satisfies both guarantees on every time interval. The Defensive Booster is very efficient: it accesses just one weak-class learner, whereas the prior online boosting methods we compare against maintain large weak-learner ensembles. Experiments on synthetic and real data streams demonstrate its strong predictive performance (sometimes substantially improving over all prior baselines) coupled with orders-of-magnitude faster runtime.