数据集蒸馏中的组合差距
The Composition Gap in Dataset Distillation
- Hokkaido University(北海道大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究数据集蒸馏中分别蒸馏的多个合成集并集的可组合性,发现即使各源精确蒸馏,组合误差仍可能显著,并推导了误差分解,表明训练保真度与下游精度是不同要求。
AI中文摘要:
数据集蒸馏将训练集压缩为一个小型合成集,通常一次评估一个。在联邦和数据治理场景中,多方各自蒸馏自己的数据,用户在其并集上训练。我们探究分别蒸馏的集合的并集是否能复现对真实数据并集进行训练的可组合性,并表明即使每个源都被精确蒸馏且总预算允许存在精确的联合蒸馏物,这种复现也可能失败。将训练轨迹压缩为更少的步骤会非线性地变换源统计量,因此对压缩后的源进行平均不同于对它们的平均进行压缩。对于二次目标,我们推导了二对一步压缩的精确组合误差,其表达式涉及源-黑森方差和损失线性项;在平滑网络上以小步长运行时,该预测在大小和方向上捕捉了局部端点差异。然而,对于学习得到的合成集,组合误差精确分解为该局部差异与一个聚合源残差之和。在端点匹配下,残差超过结构项一个数量级以上;在分布匹配下,两项部分抵消。当两个集合从同一数据集蒸馏时,联合蒸馏还保留了精度优势,此时局部差异恰好为零。因此,训练保真度和下游精度是不同的要求,两者都不能通过单独评估每个集合来确定。
英文摘要:
Dataset distillation compresses a training set into a small synthetic set, usually evaluated one at a time. In federated and data-governance settings, several parties distill their own data and a user trains on their union. We ask whether the union of separately distilled sets reproduces training on the union of the real data composability and show that it can fail even when every source is distilled exactly and the total budget admits an exact joint distillate. Compressing a training trajectory into fewer steps transforms the source statistics nonlinearly, so averaging compressed sources differs from compressing their average. For quadratic objectives we derive the exact composition error for two-to-one step compression in terms of the source-Hessian variance and the linear terms of the losses, and on a smooth network at small step sizes this prediction captures the local endpoint discrepancy in magnitude and direction. For learned synthetic sets, however, the composed error decomposes exactly into this local discrepancy and an aggregate source residual. Under endpoint matching the residual exceeds the structural term by more than an order of magnitude, and under distribution matching the two terms partly cancel. Joint distillation also retains an accuracy advantage when both sets are distilled from the same dataset, where the local discrepancy is exactly zero. Training fidelity and downstream accuracy are therefore distinct requirements, neither established by evaluating each set on its own.