发表机构
Sharif University of Technology; The University of Hong Kong(谢里夫理工大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出随机傅里叶特征高斯过程注意力模块,将注意力视为GP后验,通过低秩近似实现线性时间复杂度的不确定性校准,在保持预测精度的同时提升校准性能。
AI 中文摘要
Transformer提供了一种最先进的建模框架,但较差的校准限制了其在安全关键应用中的可靠性。一个有前景的方向是将注意力解释为高斯过程(GP)后验,这能够实现有原则的不确定性校准,但由于核的求逆,在序列长度上产生三次方复杂度;尽管解耦的GP变体将成本降低到二次方,但计算在实践中仍然令人望而却步。在本文中,我们提出了即插即用的随机傅里叶特征高斯过程注意力(RFF-GPA)模块,该模块将注意力表示为具有由随机傅里叶特征近似的平稳核的GP。这种低秩近似使得近似后验均值和方差具有线性时间复杂度,使其比先前的工作更具可扩展性。在多个真实世界数据集上的实证结果表明,我们的注意力模块在保持预测准确性的同时改善了校准,并将计算复杂度同时降低到序列长度的线性。
英文摘要
Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.
Comments14 pages, 3 figures, 3 tables