发表机构
University of São Paulo; Federal University of São Carlos(圣保罗大学; 圣卡洛斯联邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视频在严重分布偏移下测试时自适应难题,提出双重蒸馏的测试时自适应框架TADD,依赖轻量级投影适配器及互补损失适应目标域,在多个视频动作识别基准测试中优于现有TTA基线,显著提升识别准确率。
AI 中文摘要
深度学习模型在计算机视觉任务中取得了先进性能,但应用于现实场景时因分布偏移会性能严重下降。测试时自适应(TTA)试图利用目标域无标签数据在推理时动态适应测试分布。然而,适应视频等连续、时间相关数据及目标域有严重偏移时,TTA仍是难题,相关研究少。为此提出通过双重蒸馏的测试时自适应(TADD)框架,依赖轻量级投影适配器弥合域差距。适配器模块在源域预训练,再用互补损失适应目标域,包括零样本蒸馏和目标蒸馏。基于冻结的CLIP主干,该方法在推理时仅引入轻量级投影适配器作为可更新组件。在三个视频动作识别基准上评估,结果表明该方法优于现有TTA基线,在UCF-HMDB、Daily-DA和Sports-DA上相比之前方法分别提升高达+3.81%、+2.63%和+3.03%。
英文摘要
Deep learning models have achieved state-of-the-art performance in several computer vision tasks. However, they experience severe performance degradation when applied to real-world scenarios due to unanticipated distribution shifts. Test-Time Adaptation (TTA) attempts to solve this problem by using unlabeled data from the target domain to dynamically adapt to the test distribution at inference time, without access to the source data. However, TTA remains a challenging problem when adapting to continuous, temporally correlated data, such as videos, and in scenarios where the target domain contains severe domain shifts. For this reason, few works in the literature explore TTA for videos under such extreme conditions. To overcome these limitations, we propose Test-time Adaptation via Dual Distillation (TADD), an online adaptation framework that relies on a lightweight projection adapter to bridge the domain gap. The adapter module is pre-trained on the source domain and then adapted to the target using our proposed complementary losses: (i) zero-shot distillation, which encourages alignment with the domain-agnostic features from a pre-trained vision-language model (VLM); and (ii) target distillation, which retains the source domain discriminative knowledge encoded in the pre-trained adapter. Built upon a frozen CLIP backbone, our method introduces this lightweight projection adapter as the sole updatable component during inference. We conducted extensive evaluations on three well-known video action recognition benchmarks: UCF-HMDB, Daily-DA, and Sports-DA. Our experiments in the closed-set scenario demonstrate that our method consistently outperforms state-of-the-art TTA baselines. Notably, our TTA approach improves upon previous methods by up to +3.81% on UCF-HMDB, +2.63% on Daily-DA, and +3.03% on Sports-DA.