表格基础模型的注意力量化
Attention Quantization for Tabular Foundation Models
浏览论文内容
中文总结 AI 辅助
针对表格基础模型,提出FP8注意力量化策略,强调训练与测试行量化误差对齐,Triton内核实现1.7倍加速且精度无显著损失。
中文摘要 AI 辅助
随着表格基础模型(tabular foundation models)的兴起与采用,优化其推理性能成为效率研究的新兴领域。尽管这些模型在架构上与基于Transformer的大语言模型(LLMs)相似,但它们的规模和服务模式差异显著。我们表明,关注点应放在注意力计算上,而非权重或KV缓存量化——后者在LLMs中更为流行。我们针对查询(queries)、键(keys)和值(values)开发了一种量化为FP8的策略,并使用显式FP8矩阵乘法指令来加速注意力计算。我们发现,将测试行的量化误差与训练行的量化误差对齐至关重要,否则精度会急剧下降。我们的Triton内核相比常规16位内核实现了最高1.7倍的加速,并且我们证明,在TabPFN-v3和TabICLv2上,在TabArena和BeyondArena基准中均无显著的精度损失。
英文摘要
With the recent rise and adoption of tabular foundation models, optimizing their inference performance becomes an emerging field for efficiency research. While the models are architecturally similar to transformer-based large language models (LLMs), the size and serving patterns differ significantly. We show that the focus should be on the attention calculation and less on weight or KV cache quantization, which are more popular in LLMs. We develop a quantization strategy for queries, keys, and values to FP8 and use explicit FP8 matrix multiplication instructions to speed up the attention calculation. We find that it is crucial to align the quantization error in the test rows with the quantization error in the training rows, as otherwise the accuracy drops drastically. Our Triton kernel achieves a speedup up to 1.7x over regular 16-bit kernels, and we show that on TabPFN-v3 and TabICLv2 there is no relevant accuracy loss across TabArena and BeyondArena.
发表机构
- Prior Labs
机构由 AI 辅助整理,请以论文原文为准。