TTSD-FAR:用于大视频语言模型中缺失模态情感识别的带Fisher锚定恢复的测试时自蒸馏
TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs
浏览论文内容
中文总结 AI 辅助
针对大视频语言模型测试时的模态缺失问题,提出带Fisher锚定恢复的测试时自蒸馏框架,在多数据集模态缺失场景下性能优于基线方法且保持稳定。
中文摘要 AI 辅助
大视频语言模型(LVLMs)在野外多模态情感识别(ER)等多模态任务上表现出卓越性能。ER本质上是多模态任务,需要联合理解面部表情、发声、语言、生物信号和手势。然而,实际部署仍面临挑战:测试时可能出现模态缺失或噪声,部分观测相对于完整模态分布可视为分布偏移。基于熵最小化或困惑度降低的SOTA测试时自适应(TTA)方法无法迁移至自回归LVLMs,而检索增强生成(RAG)在观测模态较弱时性能下降。由于不存在验证个体更新的 ground-truth 监督,该流的自适应存在累积漂移的风险,一旦模型偏离可靠解决方案便会性能下降。因此,有效解决方案必须同时适应任意缺失模态模式并在持续自适应期间保持有效性。我们通过测试时自蒸馏(TTSD)解决这两个问题,这是一种参数高效框架,其中在完整模态上训练的冻结教师模型通过自蒸馏指导自适应低秩学生模型,仅更新极少参数。Fisher锚定恢复(FAR)将稳定性内置到同一循环中,该方法监测Fisher信息稳定性以检测收敛与漂移,当识别到分布偏移时将学生模型恢复至教师的锚点。我们在MELD、DFEW和BAH数据集上,在0%-50%的模态缺失率下进行实验,结果表明,这种统一的自适应-恢复设计在长自适应周期内始终优于基于熵的自适应、RAG和基于困惑度的生成,而无恢复机制的基线会逐渐性能下降,TTSD-FAR则保持稳定。
英文摘要
Large video-language models (LVLMs) have shown remarkable performance on multimodal tasks like multimodal emotion recognition (ER) in the wild. ER is inherently multimodal, requiring a joint understanding of facial expressions, vocalizations, language, biosignals, and gestures. However, real-world deployment remains challenging: modalities may be missing or noisy at test time. Partial observations can be viewed as a distribution shift relative to the complete-modality distribution. SOTA TTA methods based on entropy minimization or perplexity reduction do not transfer to autoregressive LVLMs, while retrieval augmented generation (RAG) degrades when the observed modality is weak. Because no ground-truth supervision exists to verify individual updates, adaptation across this stream risks accumulating drift and degrading once the model departs from a reliable solution. An effective solution must therefore adapt to arbitrary missing-modality patterns and remain effective during continual adaptation. We address both jointly with Test-Time Self-Distillation (TTSD), a parameter-efficient framework in which a frozen teacher, trained on complete modalities, guides an adaptive low-rank student via self-distillation, updating only a negligible number of parameters. Stability is built into this same loop through Fisher-Anchored Restoration (FAR), which monitors Fisher information stability to detect convergence versus drift and restores the student toward the teacher's anchor when distributional shifts are identified. Our experiments on MELD, DFEW, and BAH under 0%-50% missing modalities show that this unified adaptation-restoration design consistently outperforms entropy-based adaptation, RAG, and perplexity-based generation over long adaptation horizons, where baselines without restoration progressively degrade while TTSD-FAR remains consistent.