AI 中文总结
针对无源缺失模态的视觉识别问题,提出轻量级无源测试时适应框架TLP,通过优化对数几率提示实现性能最高提升8%。
AI 中文摘要
视觉-语言模型(Vision-Language Models, VLMs)通过利用大规模图像-文本对的互补信息,已取得了显著的性能。然而,在实际部署中,输入缺失模态的情况十分常见,往往会导致性能大幅下降。现有方法主要通过从源训练数据中学习模态补偿策略来提升模型鲁棒性,但它们依赖源训练数据,当因隐私、存储或可访问性限制无法获取原始数据时,这些方法难以应用,例如临床应用和个性化AI服务场景。这引出了一个重要但尚未充分探索的问题:能否在不访问源训练数据的情况下,于测试时高效地调整VLMs,以应对视觉识别中的缺失模态问题?为此,我们提出了测试时对数几率提示(Test-Time Logit Prompting, TLP),这是一种用于缺失模态视觉识别的轻量级无源测试时适应框架。为解决缺失模态引发的预测偏移问题,TLP通过感知不确定性调整和模态完整一致性正则化来优化对数几率提示,在保留语义一致性的同时自适应调整预测置信度。在多个不同的视觉-语言基准上开展的大量实验表明,TLP在缺失模态场景下可持续提升识别性能,最高可达8%的性能提升,且仅需数百个可调整参数和少量测试时优化步骤。
英文摘要
Vision-language models (VLMs) have achieved remarkable performance by leveraging complementary information from large-scale image-text pairs. However, missing-modality inputs are commonly encountered during real-world deployment, often leading to significant performance degradation. Existing methods primarily enhance model robustness by learning modality compensation strategies from source training data. However, their reliance on source training data makes them difficult to apply when original data are unavailable due to privacy, storage, or accessibility constraints, such as clinical applications and personalized AI services. This raises an important yet underexplored question: can VLMs be efficiently adapted at test time for visual recognition with missing modalities without accessing source training data? To this end, we propose Test-Time Logit Prompting (TLP), a lightweight source-free test-time adaptation framework for visual recognition with missing modalities. To address missing-induced prediction shifts, TLP optimizes logit prompts with uncertainty-aware adjustment and modality-complete consistency regularization, adaptively adjusting prediction confidence while preserving semantic consistency. Extensive experiments across diverse vision-language benchmarks demonstrate that TLP consistently enhances recognition performance under missing-modality scenarios, achieving up to 8\% improvements while requiring only hundreds of tunable parameters and a few test-time optimization steps.
Comments9 pages