发表机构
CUNY City College of NY; Wyze Labs, Inc.; University of Science and Technology of China; Institute of Computing Technology, Chinese Academy of Sciences; Tsinghua University(纽约城市大学城市学院; 威智实验室公司; 中国科学技术大学; 中国科学院计算技术研究所; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TAM框架,通过组织教师模型知识为可检索历史,在不增加学生推理成本的情况下,提升了视频、天气、交通流等时空预测任务的性能。
AI 中文摘要
知识蒸馏通过将精确教师模型的知识迁移至紧凑学生模型,实现高效时空预测。然而,仅针对每个样本独立匹配输出或特征,会导致跨样本预测结构未被充分利用。要利用该结构,需要能反映每个任务动态的表示与历史参考。我们提出TAM(Task-Aware Memory Distillation,任务感知知识蒸馏)框架,它将冻结教师模型的知识组织为有界、可检索的历史。记忆条目编码潜在特征、预测变化或流残差,而特定任务的选择规则可识别相关历史参考。学生模型要么匹配教师在共享参考上的相似度分布,要么回归观测条件下的残差原型。这些目标与监督预测及传统知识蒸馏形成互补。教师模型、记忆模块和辅助适配器仅在训练阶段使用,学生模型的推理过程保持不变。我们在视频预测、天气预报和交通流预测任务中,针对多种教师-学生配置对TAM进行评估。在四次配对运行的平均值中,相较于对应的KD(Knowledge Distillation,知识蒸馏)基线,加入TAM后,六个视频数据集的SSIM均得到提升,五个数据集的MSE有所降低。在KittiCaltech数据集上,平均配对MSE降低达1.86%;在使用gSTA教师模型的WeatherBench数据集上,平均配对MSE降低达1.93%;在TaxiBJ数据集上,平均配对MSE降低达1.01%。这些结果表明,在不增加学生模型推理成本的前提下,利用历史教师监督可在不同预测任务中发挥效用。
英文摘要
Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student. However, matching outputs or features independently for each sample leaves cross-sample predictive structure underused. Exploiting this structure requires representations and historical references that reflect the dynamics of each task. We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history. Memory entries encode latent features, forecast changes, or flow residuals, while task-specific selection rules identify relevant historical references. The student either matches the teacher's similarity distribution over shared references or regresses observation-conditioned residual prototypes. These objectives complement supervised prediction and conventional distillation. The teacher, memory, and auxiliary adapters are used only during training, leaving student inference unchanged. We evaluate TAM on video prediction, weather forecasting, and traffic flow prediction across multiple teacher-student configurations. Averaged over four paired runs, adding TAM improves SSIM on all six video datasets and reduces MSE on five relative to the corresponding KD baselines. Mean paired MSE reductions reach 1.86% on KittiCaltech, 1.93% on WeatherBench with a gSTA teacher, and 1.01% on TaxiBJ. These results demonstrate the utility of historical teacher supervision across distinct forecasting tasks without additional student inference cost.
Comments19 pages