适应性的代价:跨学习问题的匹配下界
The Cost of Adaptivity: Matching Lower Bounds Across Learning Problems
浏览论文内容
中文总结 AI 辅助
该研究形式化了学习问题中的干扰适应性与鲁棒性代价,推导了高斯认证的有限水平组合定律,给出匹配下界,通过实验验证了相关预测。
中文摘要 AI 辅助
自适应过程必须在没有神谕可能使用的干扰信息(如梯度尺度或光滑性指数)的情况下运行,而鲁棒过程可能需要回答那些在看到数据后才选择坐标和检查时间的查询。只有当神谕优势和有效性契约被明确陈述时,这类比较才有意义。我们通过切片归一化极小极大比来形式化干扰适应性,该比值保留每个干扰切片内的最坏情况实例,并单独定义从一个预先宣布的高斯查询扩展到任意事后检查的鲁棒性代价。我们的主要结果是高斯认证的有限水平组合定律:在M个独立坐标中,保护每个坐标和时间直至T的族级认证器,在样本均值中心化矩形类内,需付出最优归一化平方半宽度,其阶为log(eM) + log log(e^eT)。Epoch stitching给出上界;跨坐标和几何时间尺度的独立高斯块增量给出匹配下界,该下界已在几何检查点网格上成立,迫使实现的最大宽度的分位数,因此选择和停止代价会增加。两种基准机制完善了整体图景:在线凸优化中未知梯度尺度具有恒定代价,而在嵌套Hölder类上的逐点适应性代价为(log n / log log n)^(s1/(2s1+1))。作为模型监控,该定律允许分析师在任意依赖数据的时间检查M个切片指标中的任意一个:朴素固定查询带的选定覆盖率急剧下降,在M=1时降至0.30,在M≥10时降至零,而Epoch-stitched认证器将族级覆盖率保持在加性迭代对数宽度代价下。实验对这两个精确预测提出了可证伪风险,且两者均通过了检验。
英文摘要
Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose coordinate and inspection time are chosen only after the data are seen. Such comparisons are meaningful only when the oracle advantage and validity contract are stated explicitly. We formalize nuisance adaptation via a slice-normalized minimax ratio retaining the worst-case instance within each nuisance slice, and separately define the robustness cost of expanding from one preannounced Gaussian query to arbitrary post-hoc inspection. Our main result is a finite-horizon composition law for Gaussian certification: from M independent coordinates, a familywise certifier protecting every coordinate and time up to T pays optimal normalized squared half-width of order log(eM) + log log(e^eT), within the sample-mean-centered rectangular class. Epoch stitching gives the upper bound; independent Gaussian block increments across coordinates and geometric time scales give a matching lower bound, already holding on a geometric checkpoint grid, forcing quantiles of the realized maximum width so selection and stopping taxes add. Two benchmark regimes complete the picture: unknown gradient scale in online convex optimization has constant cost, while pointwise adaptation over nested Holder classes costs order (log n / log log n)^(s1/(2s1+1)). Cast as model monitoring, the law lets an analyst inspect any of M slice metrics at any data-dependent time: the naive fixed-query band's selected coverage degrades sharply, to 0.30 at M=1 and to zero for M>=10, while the epoch-stitched certifier holds familywise coverage at an additive iterated-logarithm width cost. Experiments put both sharp predictions at risk of refutation; both survive.