图注意力应何时稀疏?学习逐边Tsallis指数
When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
- University College London(伦敦大学学院)
- Holistic AI
- University of Utah(犹他大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出可学习Tsallis图注意力(LTGA),其Tsallis熵指数q可逐边学习,在8个基准测试中验证了其剪枝注意力系数的可解释机制,虽未显著优于调优的α-entmax,但可单次运行替代网格搜索。
AI中文摘要:
图注意力通过softmax归一化邻域得分,这是Shannon统计下的最大熵选择。但同配和异配图需要不同的注意力形状,单一固定归一化无法同时满足两者。我们提出LTGA(可学习Tsallis图注意力),这是一种图注意力层,其Tsallis熵指数q与权重联合学习,在从全局标量到逐边指数的四种粒度下,在重尾(q<1)、softmax(q=1)和紧支撑(q>1)注意力间连续插值,采用有界重参数化使所有模型从GAT基线开始。在8个基准测试、10次随机种子下,LTGA-Edge取得最佳平均排名(2.75),但综合检验未拒绝原假设(p=0.199),且学习q的表现不优于搜索q:验证集调优的冻结网格准确率达61.4%,调优的α-entmax为62.2%,容量匹配的q≡1对照组为62.0%,而LTGA-Edge为61.7%。学习到的指数带来的好处是单次运行而非网格搜索,以及可解释机制:当q偏离1时,它会将42%的注意力系数剪枝为精确零值,这些边是有选择性的错误边,恢复它们会损失7.1个百分点,而相同比例的随机剪枝会多损失13.0个百分点。项目页面:this https URL
英文摘要:
Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. We propose \textbf{LTGA} (\textbf{L}earnable \textbf{T}sallis \textbf{G}raph \textbf{A}ttention), a graph attention layer whose Tsallis entropic index $q$ is learned jointly with the weights, interpolating continuously between heavy-tailed ($q\!<\!1$), softmax ($q\!=\!1$) and compact-support ($q\!>\!1$) attention at four granularities from a global scalar to a per-edge index, under a bounded reparameterization that starts every model at the GAT baseline. Across eight benchmarks at ten seeds, LTGA-Edge takes the best average rank ($2.75$), but the omnibus test does not reject ($p\!=\!0.199$) and learning $q$ does not beat searching it: a validation-tuned frozen grid reaches $61.4\%$, tuned $α$-entmax $62.2\%$ and a capacity-matched $q\!\equiv\!1$ control $62.0\%$, against $61.7\%$ for LTGA-Edge. What the learned index buys is one run instead of a grid, and an interpretable mechanism: where $q$ leaves $1$, it prunes $42\%$ of attention coefficients to exactly zero, and those edges are selectively the wrong ones, restoring them costs $7.1$ points, while random pruning at the same rate costs $13.0$ more. Project page: https://kleyt0n.github.io/ltga