发表机构
National University of Singapore; The University of Tokyo(新加坡国立大学; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对联邦视频领域自适应的时间信息对齐难题,提出METAL框架,通过多尺度时间对齐与知识蒸馏实现跨域视频动作识别,在两个数据集上较现有方法最高提升28.47%,验证了多尺度设计的有效性。
AI 中文摘要
联邦视频领域自适应(FVDA)可在分布式非独立同分布(non-IID)视频数据集上开展协作学习并保护隐私,但因时间信息对齐的挑战而研究不足。我们提出多尺度时间域对齐框架(METAL),该框架利用多分辨率时间信息,仅通过模型参数迁移提升跨域视频动作识别性能。METAL在源客户端训练各尺度的Transformer编码器,随后在目标服务器对每个时间尺度执行独立知识投票以生成鲁棒伪标签;新颖的$L_2$方差惩罚项在基于尺度的知识蒸馏过程中强制跨尺度一致性,防止单一尺度主导;后期融合模块聚合不同尺度的特征,融合头通过知识蒸馏训练,利用尺度预测的置信度加权聚合,使模型能有效利用互补时间信息生成最终预测。在Epic-Kitchens-55和Daily-DA数据集上的实验表明,该方法取得了当前最优性能,相比现有FDA方法提升最高达28.47%; ablation研究证实多尺度蒸馏与尺度协调对有效时间知识迁移至关重要。
英文摘要
Federated Video Domain Adaptation (FVDA) enables collaborative learning across distributed and non-IID video datasets while preserving privacy, but is under-explored due to challenges in aligning temporal information. We propose Multi-scalE Temporal domAin aLignment (METAL), a novel framework that leverages temporal information at multiple resolutions to improve cross-domain video action recognition with only model parameter transfers. METAL trains per-scale transformer encoders on source-clients, then performs independent knowledge voting at each temporal scale to generate robust pseudo-labels on the target-server. A novel $L_2$ variance penalty enforces cross-scale consistency during scale-based knowledge distillation, preventing a singular dominant scale. The late fusion aggregates features across different scales, where the fusion head is trained via knowledge distillation using confidence-weighted aggregation of scale-wise predictions, enabling the model to effectively exploit complementary temporal information for final predictions. Experiments on Epic-Kitchens-55 and Daily-DA demonstrate state-of-the-art performances, with gains up to 28.47% over current FDA methods. Ablation studies prove that multi-scale distillation and scale coordination are critical for effective temporal knowledge transfer.
Comments10 pages, 1 figure, 5 tables. Open-source will be available