arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04424cs.LGstat.ML

稀疏注意力中信息损失与泛化之间的权衡

On the Trade-off Between Information Loss and Generalization in Sparse Attention

Zhongqi Fan, Zheng Tan

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过JS散度和Rademacher复杂度等理论工具,系统分析了稀疏注意力中信息损失与泛化性能的权衡,给出了闭式近似误差和稀疏度相关的泛化界。

中文摘要 AI 辅助

为缓解Transformer的二次复杂度瓶颈,稀疏注意力已成为一项关键技术。尽管稀疏Transformer取得了广泛的实证成功,但对其理论理解仍较为零散。特别是,两个基本问题仍不清楚:(1)稀疏化如何影响注意力机制的信息保真度?(2)这种信息损失如何与模型的泛化行为相互作用?为弥合这一差距,本文提出了对Jensen-Shannon(JS)散度以及稀疏注意力机制泛化差距的系统分析。具体而言,我们首先通过JS散度刻画近似误差。通过对截断质量α的基于阶统计量的集中分析——其中注意力分数被假设为参数为σ的独立同分布亚高斯随机变量——全注意力分布与稀疏注意力分布之间的JS散度被证明具有闭式形式log 2 + ((1 - α)/2) log(1 - α) - ((2 - α)/2) log(2 - α)。随后,我们通过Rademacher复杂度推导出一个泛化界,其量化形式为O(γ * sqrt(M/n) * (sqrt(log(3eL/M)) + sqrt(π)/2))。此外,基于互信息的泛化界以及稀疏假设类的熵和覆盖数分析,我们得到了依赖于稀疏度的泛化界O(sqrt((M/(2n)) * (log(eL/M) + log(1 + 2/ε))))。我们的分析表明,稀疏化降低了所考虑假设类的复杂度,同时引入了可通过JS散度量化的近似误差。这些发现为稀疏Transformer架构中信息保真度与泛化之间的权衡提供了理论刻画。

英文摘要

To mitigate the quadratic complexity bottleneck of the Transformer, sparse attention has emerged as a pivotal technology. Despite the extensive empirical success of sparse Transformers, the theoretical understanding of sparse attention remains fragmented. In particular, two fundamental questions remain unclear: (1) How does sparsification affect the information fidelity of attention mechanisms? (2) How does this information loss interact with the generalization behavior of the model? To bridge this gap, this paper proposes a systematic analysis of the Jensen-Shannon (JS) divergence and of the generalization gap of sparse attention mechanisms. Specifically, we first characterize the approximation error via the JS divergence. Through an order-statistics-based concentration analysis of the truncation mass alpha --- where the attention scores are assumed to be independent and identically distributed sub-Gaussian random variables with parameter sigma --- the JS divergence between the full attention distribution and the sparse attention distribution is shown to admit the closed form log 2 + ((1 - alpha)/2) log(1 - alpha) - ((2 - alpha)/2) log(2 - alpha). Subsequently, we derive a generalization bound through Rademacher complexity, quantified by O(gamma * sqrt(M/n) * (sqrt(log(3eL/M)) + sqrt(pi)/2)). Furthermore, building on a mutual-information-based generalization bound together with an entropy and covering-number analysis of the sparse hypothesis class, we obtain the sparsity-dependent generalization bound O(sqrt((M/(2n)) * (log(eL/M) + log(1 + 2/epsilon)))). Our analysis shows that sparsity reduces the complexity of the considered hypothesis class while introducing approximation error that can be quantified by the JS divergence. These findings provide a theoretical characterization of the trade-off between information fidelity and generalization in sparse Transformer architectures.

发表机构

  • Beijing Normal–Hong Kong Baptist University(北京师范大学-香港浸会大学)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑