arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PSM:基于难度精确统计匹配的数据集蒸馏

PSM: Dataset Distillation Based on Precise Statistical Matching by Difficulty

Hongxu Ma, Guang Li, Shijie Wang, Dongzhan Zhou, Suorong Yang, Baoli Sun, Takahiro Ogawa, Miki Haseyama, Zhihui Wang

arXiv 2609.34299首次发表:更新:

发表机构

Zhejiang University; Hokkaido University; The University of Queensland; Shanghai Artificial Intelligence Laboratory; National University of Singapore; Dalian University of Technology(浙江大学; 北海道大学; 昆士兰大学; 上海人工智能实验室; 新加坡国立大学; 大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有统计匹配蒸馏忽略样本难度差异的问题,提出基于难度的精确统计匹配(PSM),通过GPS划分难度组、SUA更新统计量及ISS初始化样本,在多个数据集和架构上提升下游性能。

AI 中文摘要

数据集蒸馏(DD)将大型原始数据集压缩为具有高训练效用的小型蒸馏数据集。解耦统计匹配方法在实现强性能的同时,大幅减少了蒸馏时间和内存开销。然而,这些方法通常使用从整个原始数据集估计的运行统计量来监督所有蒸馏样本。这些统计量主要捕捉平均特征分布,而忽略了样本难度之间的差异,限制了其刻画原始数据难度结构的能力。为解决此问题,我们提出了基于难度的精确统计匹配(PSM)。预训练后,PSM使用全局精确度分数(GPS)估计图像难度,对每个类别内的样本进行排序,并将每个类别划分为IPC(每类图像数)个难度组。在蒸馏过程中,统计量再次更新(SUA)通过对每个组的原始样本进行前向传播来更新教师模型的批归一化(BN)运行统计量,为对应的蒸馏批次提供特定难度的监督。同时,初始样本筛选(ISS)使用相应难度组的原始图像初始化蒸馏样本,为精确匹配提供有效的起点。在多个数据集和模型架构上的实验表明,PSM拓宽了蒸馏样本的难度范围,并在大多数评估设置中提高了下游性能。代码将发布。

英文摘要

Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distilled samples using running statistics estimated from the entire original dataset. These statistics mainly capture the average feature distribution while overlooking differences in sample difficulty, limiting their ability to characterize the difficulty structure of the original data. To address this issue, we propose Precise Statistical Matching (PSM) by difficulty. After pretraining, PSM uses the Global Precision Score (GPS) to estimate image difficulty, ranks the samples within each class, and partitions each class into IPC (images per class) difficulty groups. During distillation, Statistics Updated Again (SUA) updates the teacher's batch normalization (BN) running statistics through forward passes on original samples from each group, providing difficulty-specific supervision for the corresponding distilled batch. Meanwhile, Initial Sample Screening (ISS) initializes distilled samples using original images from the corresponding difficulty group, providing an effective starting point for precise matching. Experiments across multiple datasets and model architectures demonstrate that PSM broadens the difficulty range of distilled samples and improves downstream performance in most evaluated settings. Code will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑