arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18027cs.LGcs.CL

L1增强注意力作为一种改进的向量相似性度量

L1 Augmented Attention as an Improved Vector Similarity Metric

Kurt Godden

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对缩放点积注意力在Transformer模型中作为相似性度量的局限,提出L1增强注意力方法,通过减去特定头的L1距离改进相似性计算,经实验在WikiText 2上取得更好效果,揭示了各层几何作用及头级专业化,有效提升语言模型相似性计算。

中文摘要 AI 辅助

缩放点积注意力将方向对齐和向量大小混在一起,限制了其在Transformer模型中作为相似性度量的有效性。我们引入了L1增强注意力,这是一种简单且可计算并行化的修改,它从点积分数中减去查询和键之间学习到的特定头的L1距离。这种混合相似性捕获了互补的几何信息。点积奖励方向对齐,而L1惩罚坐标偏差。为了降低L1计算成本,我们将查询和键投影到低维子空间,其参数专门用于保留信息丰富的L1结构。在WikiText 2上使用紧凑型变压器进行评估时,L1增强注意力比原始变压器基线的困惑度降低了14.5%,并且优于RBF L2核。对范数方差和学习到的L1权重的分析揭示了各层不同的几何作用和强大的头级专业化。这些结果表明,用L1几何丰富注意力为现代语言模型中的相似性计算提供了有原则且有效的改进,对准确性和并行效率都有实际好处。

英文摘要

Scaled dot product attention conflates directional alignment and vector magnitude, limiting its effectiveness as a similarity metric in Transformer models. We introduce L1 augmented attention, a simple and computationally parallelizable modification that subtracts a learned, head specific L1 distance between queries and keys from the dot product score. This hybrid similarity captures complementary geometric information. Dot product rewards directional alignment, while L1 penalizes coordinate deviations. To reduce the cost of L1 computation, we project queries and keys into low dimensional subspaces whose parameters specialize to preserve informative L1 structure. Evaluated on WikiText 2 using a compact transformer, L1 augmented attention achieves up to a 14.5% reduction in perplexity over the original transformer baseline and outperforms an RBF L2 kernel. Analysis of norm variance and learned L1 weights reveals distinct geometric roles across layers and strong head level specialization. These results demonstrate that enriching attention with L1 geometry provides a principled and effective improvement to similarity computation in modern language models, with practical benefits for both accuracy and parallel efficiency.

补充信息

↑