arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01074cs.LGcs.CV

用于单样本测试时自适应的Logit-Origin中心化

Logit-Origin Centering for Singleton Test-Time Adaptation

Mayank Sharma, Rohit Kumar Mourya, Pratik Mazumder

首次发表
浏览论文内容

中文总结 AI 辅助

针对单样本表格全测试时自适应的退化问题,提出PLOC方法,该方法保持源模型冻结,仅存储过往logit均值,在五个表格基准上显著优于现有基线。

中文摘要 AI 辅助

表格数据广泛应用于众多实际场景,深度学习模型已被开发用于处理表格数据,但当测试数据分布与训练数据分布不同时,这些模型通常表现较差。研究人员提出了测试时自适应方法来应对该问题。全测试时自适应(FTTA)设置涉及仅使用未标记的测试数据,将部署的分类器适配到偏移的目标分布。主流FTTA方法从计算机视觉文献中继承了依赖批次的方法。本文首次证明,此类方法在严格的流式场景中会急剧退化,该场景中样本逐一到达且必须逐一分类,原因在于批次大小为1时,批次级统计量不可用或估计效果差。我们认为,单样本表格FTTA并非普通FTTA的小批量变体,而是一种独特的可识别性问题,其中仅模型的分数流位置可直接观测。为解决该问题,我们提出了轻量型方法Prequential Logit-Origin Centering(PLOC),该方法保持源模型冻结,在每一步对logit空间进行偏移。PLOC仅存储单个运行数(过往logit的均值),无需标签,无需估计先验,完全绕过权重更新。其延迟变体应用静态偏移,可精确保留源排序及AUROC。在五个表格基准、三种架构(MLP、FT-Transformer和TabTransformer)以及五个独立源检查点上进行评估,PLOC显著优于强大的表格方法和基于熵的基线。

英文摘要

Tabular data is used extensively in many real-world use cases. Deep learning models have been developed to deal with tabular data, but generally perform poorly when the test data distribution differs from that of the training data. Researchers have proposed test-time adaptation approaches to deal with this problem. The fully test-time adaptation (FTTA) setting involves adapting deployed classifiers to shifted target distributions using only unlabeled test data. Leading FTTA methods inherit a batch-dependent approach from computer vision literature. This paper demonstrates for the first time that such approaches degrade sharply in strict streaming regimes where examples arrive and must be classified one at a time. This occurs because at a batch size of one, batch-level statistics become unavailable or poorly estimated. We argue that singleton tabular FTTA is not merely a small-batch variant of ordinary FTTA, but a distinct identifiability problem where only the location of the model's score stream remains directly observable. To address this, we propose Prequential Logit-Origin Centering (PLOC), a lightweight approach that keeps the source model frozen and shifts the logit space at each step. PLOC stores only a single running number (the mean of past logits), requires no labels, estimates no priors, and bypasses weight updates entirely. A deferred variant applies a static shift that preserves the source ranking, and thus the AUROC, exactly. Evaluated across five tabular benchmarks, three architectures (MLP, FT-Transformer, and TabTransformer), and five independent source checkpoints, PLOC significantly outperforms strong tabular and entropy-based baselines.

发表机构

  • Indian Institute of Technology Jodhpur(印度焦特布尔印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑