arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IDeaL:基于改进型枯叶(Improved Dead Leaves)的无数据多教师蒸馏

IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

Feyza Yavuz, Mert Bülent Sarıyıldız, Diane Larlus

arXiv 2608.24759首次发表:更新:

发表机构

NAVER LABS Europe(NAVER欧洲实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多教师蒸馏需依赖真实数据的前提,提出IDeaL无数据蒸馏方法,通过去相关损失生成改进样本,所得学生模型性能可与基于少量真实图像蒸馏的模型媲美。

AI 中文摘要

多教师蒸馏已成为一种将互补的教师模型组合成单个学生模型的方法,该学生模型展现出所有教师的优势。学生模型被训练为在一组图像上模仿教师的输出,通常是各个教师训练集的并集,前提是这些数据可用。在本文中,我们对这一前提提出质疑并探索替代方案。我们首先研究当从输入不同类型噪声的教师模型进行蒸馏时,能达到何种程度的效果。随后,我们表明可以利用教师模型中包含的信息来为多教师蒸馏定制噪声:我们提出一种方法,得益于补丁级和图像级的去相关损失,该方法能生成针对无数据蒸馏优化的、教师特定的改进样本。实验表明,我们最有效的样本IDeaL可生成强大的学生模型,这些模型成功捕获了教师模型的互补信息,产生了极具竞争力的结果,大幅缩小了与从真实图像蒸馏得到的学生模型之间的差距。此外,在蒸馏预算为1000张图像的情况下,使用我们的IDeaL样本蒸馏得到的学生模型,其性能与使用ImageNet的1000张图像子集蒸馏得到的学生模型相当或更优。

英文摘要

Multi-teacher distillation has emerged as a way to combine complementary teacher models into a single student model that exhibits the strengths of all its teachers. The student is trained to mimic the output of the teachers on a set of images, typically the union of the individual teacher's training sets, assuming this data is available. In this paper, we question that assumption and explore alternative options. We first study how far one can go when distilling from teachers fed with different types of noise. Then, we show that information contained in the teachers can be leveraged to tailor the noise for multi-teacher distillation: we propose a method that, thanks to decorrelation losses at both patch and image levels, generates teacher-specific, improved samples optimized for data-free distillation. Experiments show that our most effective samples, IDeaL, lead to strong students that successfully capture complementary information from the teachers, yielding surprisingly competitive results that substantially narrow the gap with students distilled from real images. Moreover, given a limited budget of 1K images for distillation, students distilled using our IDeaL samples match or surpass the performance of those distilled using a 1K-image subset of ImageNet.

CommentsAccepted at ECCV 2026. Project Page is at https://blisgard.github.io/ideal_project

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑