arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10317cs.LG

深度思考:表格基础模型的回溯推断

Thinking in Depth: Retrospective Inference for Tabular Foundation Models

  • Nanjing University(南京大学)
  • National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Hao-Run Cai, Si-Yang Liu, Zi-Jian Cheng, Kun-Yang Yu, Jin-Hao Sheng, Guo Yu, Chonghan Liu, Zhi Zhou, Jun-Peng Jiang, Lan-Zhe Guo, Han-Jia Ye

AI总结:

针对表格基础模型预测细化不均的问题,提出基于回溯推断的Retro模型,通过注意力残差和查询条件门控注意力重用中间信息,在多个基准上位列前三并达到帕累托前沿。

AI中文摘要:

表格基础模型(TFMs)在多种表格任务上进行预训练,并在推理时利用新表格中的带标签示例作为上下文进行预测。最近的TFMs大多通过堆叠的Transformer层执行这种上下文预测,反复变换示例的表示方式和比较方式。通过追踪多个强TFMs中的单个查询,我们发现预测的细化在深度上高度不均匀,且往往集中在较后的层中。这种不均匀的细化促使我们重新思考中间表示在整个网络中的构建和重用方式。我们引入了Retro,一种基于回溯推断的表格基础模型,其中较后的阶段可以显式地重新访问并重组网络中较早阶段产生的中间信息。Retro围绕两个互补的操作组织这一过程:应重新访问哪些中间信息,以及如何为每个查询塑造由此产生的上下文更新。注意力残差通过自适应地重新加权来自不同深度的贡献来解决前者,而查询条件门控注意力则通过逐元素地调制表示维度上的注意力输出来解决后者。我们的分析表明,Retro将预测细化转移到更早且更广泛地跨越深度,不同阶段以暗示多视图细化的模式修订不同的查询子集。在TabArena、TALENT和RelArena上,Retro位列前三,并位于帕累托前沿。这些结果表明,直接重用中间表示为更好地利用TFMs中的深度提供了一种实用方法。

英文摘要:

Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.

↑