arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31306cs.LG

表格基础模型注意力机制基准测试

Benchmarking Attention for Tabular Foundation Models

Maximilian Schambach, Clemens Biehl, Sam Thelin

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对表格基础模型中的二维注意力机制,建立了可复现基准,评估多种注意力后端在三种GPU上的性能,发现最优选择因行列注意力及硬件而异,为表格原生注意力优化奠定基础。

中文摘要 AI 辅助

表格上下文学习模型(如TabPFN、Mitra或ConTextTab)依赖于在潜在嵌入的二维序列上交替进行行注意力和列注意力。这些注意力模式与语言模型中的一维情况显著不同:行注意力涉及更长的序列,而列注意力则处理更短的序列,且表格数据的跨步内存布局使得生成连续张量的成本高昂。此外,当前模型中使用的隐藏维度相比近期语言模型较小。然而,高效注意力机制的研究主要集中在一维序列上,使得二维表格场景未被探索。为此,我们创建了一个可复现的基准测试设置,并研究了表格注意力在多种后端——Torch SDPA(高效和cuDNN)、FlashAttention-2/3/4,以及仅推理后端vLLM和SageAttention——中的独特特征,在三种GPU代际(A100、H100、B200)上测量了实际表格形状的前向和后向吞吐量。我们发现,最优后端选择在列注意力和行注意力之间有所不同,并随硬件及模型特性而变化:虽然针对各GPU代际定制的FlashAttention实现整体表现最佳,但在较长序列的列注意力情况下,它们有时会被CuDNN超越,交叉点取决于头维度。在仅推理后端中,SageAttention在行注意力和超过16k行的大序列上表现良好。我们的可复现基准为未来表格原生注意力的改进奠定了基础。自包含的基准测试和评估代码可在以下网址公开获取:this https URL

英文摘要

Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16\,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark

发表机构

  • SAP SE(SAP公司)

机构由 AI 辅助整理,请以论文原文为准。

↑