arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06679cs.LGcs.AI

弥合上下文差距:面向表格上下文学习的激活对齐

Closing the Context Gap: Activation Alignment for Tabular In-Context Learning

Yoel Zeldes

首次发表
浏览论文内容

中文总结 AI 辅助

提出激活对齐方法,通过线性变换将部分上下文学生模型的激活映射至全上下文教师模型,在不增加推理成本下显著提升表格上下文学习性能,恢复近半性能差距。

中文摘要 AI 辅助

表格基础模型通过将预测条件建立在作为上下文提供的带标签训练示例上,执行上下文学习(ICL)。与传统将训练与推理分离的模型不同,这些模型必须在每次前向传播中处理所有训练示例,使得每次预测成本高昂。限制训练示例的数量可降低此成本,但会显著降低性能。我们不丢弃上下文,而是提出激活对齐方法,该方法利用完整上下文来教导模型在仅看到子集时如何表现。这是通过在合成无标签数据上训练一个轻量级线性变换实现的,将数据受限的“学生”(使用部分上下文)的中间激活映射到使用全部数据的全上下文“教师”的激活。训练对齐器无需GPU,在普通硬件上可在数秒至数分钟内收敛。我们使用领先的两个表格基础模型TabPFN-3和TabFM,在TabArena基准的38个分类数据集上进行了评估。在所有上下文预算下,对齐后的学生模型相对于未对齐的基线,对两个模型均产生了广泛且统计显著的改进。在低数据场景下,对齐恢复了教师预测优势的近一半。该方法提供了一种实用、低开销的方式,在实现紧凑上下文推理速度的同时,弥合了与全上下文教师性能差距的显著部分。

英文摘要

Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we propose activation alignment, a method that leverages the full context to teach a model how to behave when seeing only a subset. This is achieved by training a lightweight linear transformation on synthetic unlabeled data to map the intermediate activations of a data-constrained "student" (using partial context) toward those of a full-context "teacher" (using all data). Training the aligner requires no GPU and converges in seconds to minutes on commodity hardware. We evaluate on 38 classification datasets from the TabArena benchmark using the leading two tabular foundation models, TabPFN-3 and TabFM. Across all context budgets, the aligned student yields broad, statistically significant improvements over the unaligned baseline for both models. In low-data regimes, alignment recovers nearly half of the teacher's predictive advantage. The method provides a practical, low-overhead approach to achieving the inference speed of compact contexts while closing a significant fraction of the performance gap to the full-context teacher.

补充信息

↑