用于细粒度视频表情识别的基于能量缓存的在线个性化测试时自适应
Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition
浏览论文内容
中文总结 AI 辅助
针对视频FER的测试时自适应难题,提出EB-CaP方法,通过轻量级能量模型结合CLIP生成个性化缓存,在低开销下优于现有SOTA方法。
中文摘要 AI 辅助
视频中的面部表情识别(FER)具有挑战性,因为模型必须识别随时间演变且在个体间存在差异的微妙情感状态。尽管视觉-语言模型提供了可迁移的视觉-语义表征,但在与主体无关的数据上训练的模型,在推理阶段遇到与主体相关的分布偏移时性能会下降。现有的测试时自适应(TTA)方法通常会在推理阶段更新模型参数,这会增加计算成本和延迟。基于缓存的方法避免了参数更新,但通常需要足够的目标样本以形成可靠的类别原型,这在适应初期以及对于罕见观察类别而言是难以实现的。我们提出了基于能量的缓存个性化(EB-CaP),这是一种用于视频FER的基于主体的在线TTA方法,可针对每个目标视频生成个性化的特定类别原型。EB-CaP使用轻量级基于能量的模型,从当前未标记的视频中采样原型并在线填充个性化缓存,无需积累大量目标数据或存储多样化的源原型。其能量函数仅依赖于预训练的CLIP:目标视频嵌入与类别文本嵌入之间的相似性指导原型采样。同时,正缓存和负缓存分别存储可靠和不确定的目标嵌入。自适应熵门根据不断变化的置信度分布控制缓存更新,而多样性门则限制冗余样本。最终预测将缓存衍生的分数与当前CLIP分数相结合。在BioVid、StressID和BAH数据集上的实验表明,EB-CaP的性能优于最先进的TTA方法,同时保持了较低的计算和内存开销。代码可在this https URL获取。
英文摘要
Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across individuals. Although vision-language models provide transferable visual-semantic representations, models trained on subject-independent data often degrade under subject-specific distribution shifts at inference time. Existing test-time adaptation (TTA) methods commonly update model parameters during inference, increasing computational cost and latency. Cache-based methods avoid parameter updates, but they usually require enough target samples to form reliable class prototypes, which is difficult early in adaptation and for rarely observed classes. We introduce Energy-Based Cache Personalization (EB-CaP), a subject-based online TTA method for video FER that generates class-specific prototypes personalized to each target video. EB-CaP uses a lightweight energy-based model to sample prototypes from the current unlabeled video and populate a personalized cache online, without accumulating large amounts of target data or storing diverse source prototypes. Its energy function relies only on pretrained CLIP: similarities between the target video embedding and class text embeddings guide prototype sampling. In parallel, positive and negative caches store reliable and uncertain target embeddings. An adaptive entropy gate controls cache updates according to the evolving confidence distribution, while a diversity gate limits redundant samples. Final predictions combine cache-derived scores with the current CLIP scores. Experiments on BioVid, StressID, and BAH show that EB-CaP outperforms state-of-the-art TTA methods while maintaining low computational and memory overhead. Code is available at https://github.com/MasoumehSharafi/EB-CaP.