发表机构
SinoPac Holdings(永丰金控)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RepICL,一种元训练的上下文预测器,通过情节白化跨异构表示空间复用少样本预测过程,在RepShiftBench的12个设置中超越逻辑回归,证明共享预测过程可泛化至未见表示空间。
AI 中文摘要
冻结的表示被广泛复用于下游分类,但每个新任务通常需要拟合新的预测器。我们探究少样本预测过程本身能否被学习一次,并在数据集和表示空间之间复用。为研究此问题,我们引入了RepShiftBench,包含文本、图像和音频上的1,218个编码器-数据集任务,并分别评估对未见数据集、未见编码器、联合未见数据集和编码器以及未见模态的泛化。该基准揭示了一个显著差距:独立拟合于每个情节的逻辑回归在所有设置中均优于所有评估的上下文学习器。我们提出RepICL,一种元训练的上下文学习器,在预测前通过情节白化对每个情节进行规范化。其归纳变体RepICL-I在所有12个基准设置中超越了逻辑回归,而RepICL-T则大幅优于现有的转导方法。消融实验确定情节白化是这些增益的主要来源,同时表明它并非普遍有益的预处理步骤。在两种变体中,增益集中于那些简单支持原型指向错误类别或对真实类别与竞争类别之间区分度低的查询。当有限的支持覆盖导致对类别分离的误导性视图时,转导提供了最大的额外增益。这些结果共同表明,共享的少样本预测过程可以泛化到训练期间观察到的表示空间之外。
英文摘要
Frozen representations are widely reused for downstream classification, yet each new task typically requires fitting a new predictor. We ask whether the few-shot prediction procedure itself can instead be learned once and reused across datasets and representation spaces. To study this question, we introduce RepShiftBench, comprising 1,218 encoder--dataset tasks across text, image, and audio, with separate evaluation of generalization to unseen datasets, unseen encoders, jointly unseen datasets and encoders, and unseen modalities. The benchmark exposes a substantial gap: Logistic Regression fitted independently on each episode outperforms every evaluated in-context learner across all settings. We introduce RepICL, a meta-trained in-context learner that canonicalizes each episode through episodic whitening before prediction. Its inductive variant, RepICL-I, surpasses Logistic Regression in all 12 benchmark settings, while RepICL-T substantially outperforms existing transductive methods. Ablations identify episodic whitening as the primary source of these gains, while showing that it is not a universally beneficial preprocessing step. Across both variants, the gains concentrate on queries for which simple support prototypes favor the wrong class or provide little separation between the true class and competing classes. Transduction provides its largest additional gains when limited support coverage gives a misleading view of class separation. Together, these results demonstrate that a shared few-shot prediction procedure can generalize beyond the representation spaces observed during training.