arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

适用于恰当损失的交换不可知学习的快速收敛速率

Fast Rates for Swap-Agnostic Learning of Proper Losses

Princewill Okoroafor

arXiv 2607.28856首次发表:更新:

发表机构

Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对恰当损失的交换不可知学习,提出了联合控制预测水平比较的方法,得到了紧的快速收敛速率,优于现有结果,核心是将该学习归约为二阶多校准问题。

AI 中文摘要

交换不可知学习强化了经典不可知学习,允许比较器在学习者预测的每个水平集上选择不同的假设。该基准捕捉了依赖预测的后处理,但似乎需要为每个可能的预测值求解单独的不可知学习问题。我们证明,对于恰当损失,这些预测水平的比较可被联合控制。我们的主要结果是针对任意固定恰当损失的离线交换不可知学习器。对于有限假设类H和任意固定光滑恰当损失,m个独立同分布样本的超额风险为\\(\widetilde{O}((\log |H|/m)^{2/3})\\),对应的在线交换遗憾界为\\(\widetilde{O}(T^{1/3}(\log |H|)^{2/3})\\)。我们还提供了算法,其预测对整个损失族同时具有交换不可知性。对于所有在[-1,1]上有界的恰当损失,我们分别得到在线速率\\(\widetilde{O}(\sqrt{T\log |H|})\\)和离线速率\\(\widetilde{O}(\sqrt{\log |H|/m})\\)。对于凸的1-Lipschitz恰当损失,这些速率提升为在线\\(\widetilde{O}(T^{1/3}(\log |H|)^{2/3})\\)和离线\\(\widetilde{O}((\log |H|/m)^{2/3})\\)。这些界在对数因子内是紧的,优于Luo等人(2025)提出的交换全预测保证隐含的\\(\widetilde{O}(T^{2/3}(\log |H|)^{1/3})\\)速率。我们的主要技术贡献是通过带有Bernstein型方差校正的Blackwell可达性,将交换不可知学习归约为二阶多校准问题。

英文摘要

Swap-agnostic learning strengthens classical agnostic learning by allowing the comparator to select a different hypothesis on each level set of the learner's predictions. This benchmark captures prediction-dependent postprocessing, but appears to require solving a separate agnostic-learning problem for every possible prediction value. We show that, for proper losses, these prediction-level comparisons can instead be controlled jointly. Our main result is an offline swap-agnostic learner for any fixed proper loss. For a finite hypothesis class $H$ and any fixed smooth proper loss, the excess risk from $m$ i.i.d. samples is $\widetilde{O}((\log |H|/m)^{2/3})$, with a corresponding online swap-regret bound of $\widetilde{O}(T^{1/3}(\log |H|)^{2/3})$. We also give algorithms whose predictions are simultaneously swap-agnostic for entire families of losses. For all proper losses bounded in $[-1,1]$, we obtain online and offline rates of $\widetilde{O}(\sqrt{T\log |H|})$ and $\widetilde{O}(\sqrt{\log |H|/m})$, respectively. For convex, $1$-Lipschitz proper losses, these rates improve to $\widetilde{O}(T^{1/3}(\log |H|)^{2/3})$ online and $\widetilde{O}((\log |H|/m)^{2/3})$ offline. These bounds are tight up to logarithmic factors and improve upon the $\widetilde{O}(T^{2/3}(\log |H|)^{1/3})$ rate implied by the swap-omniprediction guarantee of Luo et al. (2025). Our main technical contribution is a reduction from swap-agnostic learning to a second-order form of multicalibration, obtained via Blackwell approachability with a Bernstein-style variance correction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑