用于复杂状态传播的基于马氏距离的多头注意力
Mahalanobis-Based Multi-Head Attention for Complex State Propagation
- GuangDong Police College(广东警官学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于马氏距离的多头注意力MHA-CSP,通过马氏距离RBF核与跨头协作实现高效结构化推理,在长序列状态跟踪任务上优于Transformer和GCN基线。
AI中文摘要:
本文提出了一种新型注意力机制——基于马氏距离的多头注意力(MHA-CSP),该机制用基于马氏距离的RBF核替代标准点积,可在不增加参数数量的情况下有效计算无限维特征空间中的注意力。关键在于,马氏距离的正定性支持直接构建树注意力:注意力分数直接由累积距离构建,并通过LogSumExp校正项,即减去边指数的对数和,来修正原始距离。此外,多头马氏距离矩阵可被重新用于构建注意力网格机制,实现跨头核协作,同时提升准确率和训练效率。大量实验表明,仅含11.9万参数且仅在最终隐藏状态应用教师强制的MHA-CSP,在长序列状态跟踪任务上,与相同条件下从头训练的Transformer和GCN基线相比,始终表现更优。这些基线依赖密集注意力或图传播,而MHA-CSP通过基于马氏距离的注意力实现的合成距离校正,以及继承自CSP骨干的高效信息旁路,实现了稳健的结构化推理。该结果凸显了具有协作多头校正的复数值状态传播在捕获符号结构方面的有效性,为结构化推理建立了新的效率-性能权衡。
英文摘要:
In this paper, we propose \textbf{Mahalanobis-Based Multi-Head Attention} (MHA-CSP), a novel attention mechanism that replaces the standard dot-product with a \textbf{Mahalanobis distance-based RBF kernel}, which effectively computes attention in an infinite-dimensional feature space without increasing the parameter count. Crucially, the positive definiteness of the Mahalanobis distance enables a \textbf{direct construction of Tree Attention}: attention scores are built directly from accumulated distances, with a LogSumExp correction that rectifies the raw distance by subtracting the log-sum of edge exponentials. Moreover, the multi-head Mahalanobis distance matrices are themselves repurposed to construct an \textbf{attention meshing mechanism}, enabling cross-head kernel collaboration that simultaneously boosts accuracy and training efficiency. Extensive experiments demonstrate that MHA-CSP, with only 119K parameters and \textbf{teacher forcing applied exclusively at the final hidden state}, consistently outperforms Transformer and GCN baselines trained from scratch under identical conditions on long-sequence state tracking tasks. While these baselines rely on dense attention or graph propagation, MHA-CSP achieves robust structured reasoning via synthetic distance rectification---powered by Mahalanobis-based attention---and efficient information bypass inherited from the CSP backbone. This result highlights the effectiveness of complex-valued state propagation with collaborative multi-head rectification in capturing symbolic structures, establishing a new efficiency-performance trade-off for structured reasoning.