通过候选验证和散度偏移实现高效的测试时自适应
Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts
浏览论文内容
中文总结 AI 辅助
本文提出测试时校正(TTC)框架,将VLM测试时自适应重构为候选验证,通过假设-重构-校正原则选择最小散度偏移的候选,实现无需训练的高效校正,在15个基准上提升精度并显著降低计算开销。
中文摘要 AI 辅助
视觉语言模型(VLMs)在零样本迁移方面表现出色,但在推理时仍易受目标域偏移的影响。测试时自适应(TTA)提供了一种实用的补救措施,然而大多数现有的VLM-TTA方法遵循预测侧自适应范式。它们利用测试样本来调整logits、原型、缓存、先验或特征统计,往往会产生额外的计算开销。在本文中,我们采取不同的视角,将VLM-TTA重新定义为候选验证而非预测调整。我们提出测试时校正(TTC),一种基于假设的校正框架,遵循一个简单的原则:假设、重构、校正。给定一个测试特征及其top-k候选标签,TTC将每个候选标签视为一个假设,在存储在记忆库中的相应潜在子空间内重构特征,并测量由此产生的散度偏移。该偏移量化了在假设插入测试特征后,候选子空间及其与其他候选的关系的变化程度。正确的候选假设仅引起较小的偏移,而错误的假设则会更强地扰动子空间。因此,TTC通过选择具有最小聚合散度偏移的候选来校正预测。这种无需训练的候选验证机制避免了迭代优化,并提供了良好的精度-效率权衡。在五种TTA设置和15个基准数据集上,包括零样本分类、域泛化、少样本分类、基础到新颖泛化和跨数据集评估,TTC在准确性上持续优于最先进的VLM-TTA方法,同时实现了高达2倍的加速、比最低内存的无需训练基线低3倍以上的CPU内存使用,以及低1.4倍的GPU内存使用。
英文摘要
Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose Test-Time Correction (TTC), a hypothesis-based correction framework guided by a simple principle: hypothesize, reconstruct, correct. Given a test feature and its top-k candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting divergence shift. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This training-free candidate-verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to 2x speedup, over 3x lower CPU memory usage, and up to 1.4x lower GPU memory usage than the lowest-memory training-free baseline.
发表机构
- Ajou University(亚洲大学)
机构由 AI 辅助整理,请以论文原文为准。