arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

损失差条件互信息的精度-信息权衡

An Accuracy-Information Tradeoff for Loss-Difference Conditional Mutual Information

Hazar Yueksel

arXiv 2610.09206首次发表:更新:

AI 中文总结

本文证明精度迫使损失差条件互信息(ld-CMI)下界,在线性预测器与光滑凸损失下,即使泛化差距为$O(n^{-1/2})$,ld-CMI仍为$n$比特量级,揭示精度与信息权衡的本质。

AI 中文摘要

损失差条件互信息(ld-CMI)使用泛化界超样本层级中标准观测的最小值:它衡量学习者的损失差揭示了关于每对候选者中训练了哪一个的信息。已知精度迫使信息进入模型;数据处理不会将这种下界传递到损失。我们通过界定损失差的三个矩,表明精度也迫使ld-CMI。对于具有非零斜率光滑凸损失(如逻辑损失)的线性预测器,加上曲率和增长均为幂$r\ge2$的正则化器,在维度至少为$n$的线性缩放的符号立方体上的乘积分布上,每个在最优样本量$n\asymp\varepsilon^{-2+2/r}$下在这些分布上期望超额风险至多$\varepsilon$的恰当学习器具有最坏情况下的ld-CMI为$n$比特量级,并且在损失差上具有标准差$\tau$的高斯噪声下为$\Theta(n/(1+(\tau/\varepsilon)^2))$比特。在没有正则化器的情况下,在$n\asymp\varepsilon^{-2}$时同样成立。因此,范围缩放的ld-CMI界不能在这些分布上消失,尽管每个恰当学习器的泛化差距为$O(n^{-1/2})$。我们还表明,模型级信息不决定有噪声的损失差信息,并且增长、斜率和维度条件是必需的,最后一个条件直到对数因子。

英文摘要

Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power $r\ge2$, on product distributions over a scaled sign cube in dimension at least linear in $n$, every proper learner with expected excess risk at most $\varepsilon$ on these distributions at the optimal sample size $n\asymp\varepsilon^{-2+2/r}$ has worst-case ld-CMI of order $n$ bits, and $Θ(n/(1+(τ/\varepsilon)^2))$ bits under Gaussian noise of standard deviation $τ$ on the loss differences. The same holds without a regularizer, at $n\asymp\varepsilon^{-2}$. Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner's generalization gap is $O(n^{-1/2})$. We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.

Commentsv2: title metadata corrected (no change to the paper). 61 pages, of which 8 pages main text. Code, data and the Lean 4 formalization are in the ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑