arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当更少计算带来更多收益:自适应早退提升预训练异常检测

When Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier Detection

Tianyang Zhou, Leman Akoglu

arXiv 2609.32898首次发表:更新:

AI 中文总结

本研究提出深度自适应早退方法,在预训练异常检测模型中通过最优中间层退出,平均提升性能4.7-7.3%,并预训练路由器实现高达1.8倍加速。

AI 中文摘要

预训练表格基础模型以固定深度处理每个数据集,推理成本随数据集规模增长。为解决这一问题,我们首次研究了面向预训练异常检测模型的深度自适应早退方法。虽然早退通常以效率为动机,但我们发现了一个令人惊讶的好处:在最优中间层退出,也能在多样化的真实世界基准上平均提升检测性能4.7-7.3%,且这一提升在三个不同的基础模型上保持一致。首先,我们调查了驱动这些收益的因素,并识别出一个关键机制:上下文污染,即上下文样本中存在异常值。我们的分析表明,邻近的上下文样本在更深层对查询预测的影响越来越大,这与这些模型的基于检索的观点一致。实际上,早退减轻了检索准确但受污染邻居的不利影响,当上下文污染与自然异常率匹配时,收益为13-21%。受这些发现的启发,我们预训练了一个即插即用的路由器,以选择数据集特定的退出层,利用查询异常标签作为仅在路由器训练期间可用的特权信息。路由器事后运行,保持基础模型参数和预测头不变。在三个大型真实世界基准上的实验表明,在干净上下文中,路由器在三个预训练骨干上恢复了高达45%的预言机收益,并实现了高达1.8倍的加速,随着上下文污染的增加,收益更大。

英文摘要

Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is typically motivated by efficiency, we uncover a surprising benefit: exiting at the optimal intermediate layer can also improve detection performance on diverse real-world benchmarks by 4.7-7.3% on average, consistent across three distinct foundation models. First, we investigate the factors driving these gains, and identify a key mechanism: context pollution, i.e., the presence of outliers among in-context samples. Our analysis reveals that nearby in-context samples exert increasing influence on query predictions at greater depths, consistent with a retrieval-based view of these models. In effect, early-exit alleviates the adverse effects of retrieving accurate-yet-polluted neighbors, with gains of 13-21% when context pollution matches the natural outlier rate. Motivated by these findings, we pretrain a plug-in router to select a dataset-specific exit layer, using query outlier labels as privileged information available only during router training. The router operates post hoc, leaving the base model parameters and prediction head unchanged. Experiments on three large real-world benchmarks show that, on clean context, the router recovers up to 45% of the oracle gain with up to 1.8x speedup across three pretrained backbones, with larger gains as context pollution increases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑