arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14020cs.LGstat.ML

更多正确数据何时会造成损害?插入稳定性与基于维度理论的局限

When Does More Correct Data Hurt? Insertion-Stability and the Limits of Dimension-Based Theory

Joseph Sankoorikal Johny

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探讨了添加正确数据可能损害学习性能的问题,定义插入稳定学习器可免疫于该问题,证明交闭类别具插入稳定性,且学习器与类别的适配性是关键而非单纯的类别或学习器优劣。

中文摘要 AI 辅助

添加已知为正确的数据本应是安全的,但并非总是如此。Larsen、Pabbaraju 和 Shetty 用单调对手对这种失败进行建模:该对手读取一个独立同分布的训练样本,并可追加任意多的额外示例,只要目标假设为它们全部打上标签。此后 Mehrotra 确定了该代价,表明对于 VC 维 d ≥ 2 的类别,没有任何学习器能保证期望误差优于 Θ((d/n)log(en/d)),这比干净 PAC 速率高出一个对数因子。由于该速率是所有类别的最坏情况,它未说明哪些类别实际会遭受该惩罚,而答案取决于学习器。我们定义:若向学习器提供更多正确标记的示例只会缩小其错误区域,则称该学习器是插入稳定的。此类学习器可免疫于该对手,因为在任意给定样本上,插入后的风险绝不会超过仅在干净部分的风险,无论添加多少数据及如何巧妙选择。高概率保证保持不变,且由于 Closure 是插入稳定的,每个交闭类别都保持其干净的 E[Err] ≤ (21d+34)/n 的速率。免疫性无法由经典维度预测:两个类别若 VCdim = Ldim = 2,仍可能存在差异,一个的速率为 Θ(1/n),另一个为 Θ(log(en)/n),而区间的 Littlestone 维度无界但仍具免疫性。在 Mehrotra 的困难类别上,我们不仅证明单个算法会失败,还表明任何有限大小的单调置换不变压缩方案都无法达到干净速率。因此问题不在于类别是否困难,也不在于学习器是否优秀,而在于二者是否适配:给定在干净数据上最优的插入稳定学习器,正确添加数据是无代价的;若无此类学习器,则代价属于类别,改变学习器无法避免该代价。

英文摘要

Adding data known to be correct ought to be safe. Not always. Larsen, Pabbaraju and Shetty model the failure with a monotone adversary, which reads an i.i.d. training sample and may append as many further examples as it likes, provided the target hypothesis labels them all. Mehrotra has since settled the cost, showing that for classes of VC dimension d >= 2 no learner can guarantee expected error better than Theta((d/n)log(en/d)), a logarithmic factor above the clean PAC rate. Because that rate is a worst case over all classes, it says nothing about which classes actually suffer the penalty, and the answer turns on the learner. We call a learner insertion-stable if feeding it more correctly labeled examples can only shrink the region where it errs. Such learners are immune to the adversary, since on any given sample the risk after insertions never exceeds the risk on the clean part alone, however much is added and however cleverly it is chosen. High- probability guarantees carry over unchanged, and because Closure is insertion-stable every intersection-closed class keeps its clean rate of E[Err] <= (21d+34)/n. Immunity is not something the classical dimensions can predict. Two classes can agree on VCdim = Ldim = 2 and still split, one at Theta(1/n) and the other at Theta(log(en)/n), while intervals have unbounded Littlestone dimension and are immune anyway. On Mehrotra's hard class we prove more than the failure of a single algorithm, showing that no monotone permutation-invariant compression scheme of any finite size attains the clean rate. The question is therefore not whether a class is hard, nor whether a learner is good, but whether the two suit each other. Given an insertion-stable learner that is optimal on clean data, correct additions are free, and without one the cost belongs to the class, so changing the learner will not avoid it.

补充信息

↑