发表机构
ETH Zürich; ELLIS Institute Tübingen; MPI-IS; ETH AI Center(苏黎世联邦理工学院; ELLIS图宾根研究所; 马克斯·普朗克智能系统研究所; ETH人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究重新审视因果训练设置,发现DeltaNet线性序列混合器在表格上下文学习中最有前景,但长度泛化受限;通过时间依赖衰减调度干预稳定循环状态,并在OpenML-CC18和TabArena上匹配softmax注意力基线。
AI 中文摘要
表格基础模型通过在上下文中对带标签的示例进行条件化处理来实现强性能,但softmax注意力机制限制了其在大型数据集上的使用。现有的线性时间替代方案大多为因果性的,其用于表格上下文学习(ICL)的潜力尚未得到充分探索。为解决这一问题,我们(1)重新审视因果训练设置,(2)比较线性序列混合器,(3)研究它们在预训练上下文长度之外的ICL泛化能力。首先,我们表明因果模型的最佳训练设置类似于下一个词元预测。然后,令人惊讶的是,最有前景的线性序列混合器是因果性的:DeltaNet甚至优于非因果线性注意力。然而,在超过预训练上下文长度2-4倍时,其性能会下降,而现有的缓解策略(如双向性)至多只能推迟问题。一个隐藏状态预言机表明这不是容量问题。相反,我们的分析指向循环状态中的不稳定性,该状态在因果模型的更深层中发生漂移。由于DeltaNet学习到的写入速率对预训练机制过拟合,我们通过时间依赖的衰减调度干预来调节它们,以稳定长度泛化。最后,通过从最终状态读出重新引入非因果性,我们能够在OpenML-CC18和TabArena上紧密匹配受控的softmax注意力基线。
英文摘要
Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large datasets. Existing linear-time alternatives, however, are mostly causal, and their potential for tabular in-context learning (ICL) remains underexplored. To address this, we (1) revisit causal training setups, (2) compare linear sequence mixers, and (3) investigate their ICL generalisation beyond the pretraining context length. First, we show that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. However, it degrades beyond $2$-$4\times$ the pretraining context length, and existing mitigation strategies such as bidirectionality defer the problem at best. A hidden-state oracle shows that this is not a capacity problem. Instead, our analysis points to an instability in the recurrent state, which drifts in deeper layers of causal models. Since DeltaNet's learned write rates overfit to the pretraining regime, we modulate them with a time-dependent decay schedule intervention to stabilise length generalisation. Finally, re-introducing non-causality by reading out from the final state allows us to closely match a controlled softmax attention baseline on OpenML-CC18 and TabArena.
Comments31 pages, 13 figures, 10 tables. Code available at https://github.com/schnurrd/ICL-Architectures