基于簇引导原型混合的高效视频数据集蒸馏
Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
浏览论文内容
中文总结 AI 辅助
针对现有视频数据集蒸馏方法迭代优化成本高的问题,提出ProtoBlend框架,通过选择-分配-混合的方式实现高效视频数据集蒸馏,在四个动作识别基准上取得了有竞争力的准确率-效率权衡。
中文摘要 AI 辅助
视频数据集蒸馏旨在将大型视频数据集压缩为紧凑的替代集,同时保留其训练效用。现有大多数方法通过迭代优化合成浓缩视频,其计算成本会因时间维度而放大。我们未进一步减少优化变量数量,而是研究能否在不对存储视频进行基于梯度的优化的情况下构建有效的蒸馏视频。这种基于构建的方法需应对三个挑战:选择信息丰富的时间片段、在每类有限视频预算下覆盖类内多样变化、提升每个存储样本携带的信息。为此,我们提出ProtoBlend,一个高效的选择-分配-混合框架:首先,教师引导的时间片段选择从每个源视频中保留高置信度片段;其次,簇引导的原型分配在教师特征空间中划分所选片段,为每个类内簇分配一个蒸馏槽位;最后,每个原型与簇内锚点混合,同时用相同系数组合其教师预测以提供混合源监督。在四个修剪后的动作识别基准上的实验表明,ProtoBlend无需对蒸馏视频进行迭代优化,即可实现具有竞争力的准确率-效率权衡。
英文摘要
Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most existing approaches synthesize condensed videos through iterative optimization, whose cost is amplified by the temporal dimension. Rather than further reducing the number of optimized variables, we investigate whether effective distilled videos can be constructed without gradient-based optimization of the stored videos. Such a construction-based approach must address three challenges: selecting informative temporal segments, covering diverse intra-class variations under a limited videos-per-class budget, and increasing the information carried by each stored sample. To this end, we propose ProtoBlend, an efficient select-allocate-blend framework. First, teacher-guided temporal clip selection retains a high-confidence segment from each source video. Second, cluster-guided prototype allocation partitions the selected clips in the teacher feature space and assigns one distilled slot to each intra-class cluster. Third, each prototype is blended with an in-cluster anchor, while their teacher predictions are combined using the same coefficient to provide mixture-source supervision. Experiments on four trimmed action-recognition benchmarks demonstrate that ProtoBlend achieves a competitive accuracy-efficiency trade-off without iterative optimization of the distilled videos.