arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于在Transformer中学习核注意力的库仑粒子模型

A Coulomb Particle Model for Learning Kernel Attention in Transformers

Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran

arXiv 2607.23869首次发表:更新:

发表机构

eBay Inc.(易贝公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于粒子方法,通过优化核目标对齐学习特征分布,用于Transformer注意力。先学习正随机特征映射,再冻结核训练其余参数。实验表明,此方法能提升多种特征映射的准确性、校准和鲁棒性,且保持线性注意力推理复杂度。

AI 中文摘要

随机特征为核机器提供了可扩展的近似,但性能很大程度上取决于特征分布的选择。我们提出一种基于粒子的方法,通过优化核目标对齐来学习这种分布,同时用里斯/库仑排斥势对粒子进行正则化。由此产生的哈密顿量产生多样的、任务自适应的随机特征,并通过麦克凯恩-弗拉索夫方程允许平均场描述。我们在线性化Transformer注意力中实例化该方法,先在第一阶段学习正随机特征映射,然后冻结核并用交叉熵训练其余网络参数。合成分类和句子级基准实验表明,学习到的核注意力可以提高多种特征映射的准确性、校准和鲁棒性,同时保持线性注意力推理复杂度。

英文摘要

Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coulomb repulsive potential. The resulting Hamiltonian yields diverse, task-adaptive random features and admits a mean-field description through a McKean--Vlasov equation. We instantiate the method in linearized Transformer attention by learning positive random-feature maps in a first alignment phase, then freezing the kernel and training the remaining network parameters with cross-entropy. Experiments on synthetic classification and sentence-level benchmarks show that learned kernelized attention can improve accuracy, calibration, and robustness for several feature maps while preserving linear-attention inference complexity.

CommentsWorkshop on High-dimensional Learning Dynamics (HiLD), ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑