arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度归一化注意力的可辨识性与可观测性

The Identifiability and Observability of Deep Normalized Attention

Pranav Venkata Konda

arXiv 2610.09620首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究了深度归一化注意力的参数可辨识性,证明了实解析情形下的猜想,并分析了坍缩导致的不可观测性及局部辨识的泰勒阶数。

AI 中文摘要

我们研究了深度、无掩码、单头注意力的哪些参数由其输入-输出函数决定。对于已知的正非恒定实解析归一化器,该函数在由偶归一化器引起的符号意义下,一般性地决定了有效分数和组合值映射。这证明了Henry-Marchietti-Kohn猜想在实解析情形下的成立,包括softmax情形。然后,我们在显式归一化器条件下对异常纤维进行分类,识别出当坍缩使得后续分数不可观测的情况,并为局部辨识建立了尖锐的泰勒阶数。在同时查询/键坍缩附近,我们在分离的有限输入库上计算了完整的本征雅可比衰减谱。对于常见的首个非恒定归一化器次数$k$,第$i$层具有接触阶$2k3^{i-1}-1$,并具有精确的重数和核维数。高精度和自动微分计算说明了由此导致的数值灵敏度损失。

英文摘要

We study which parameters of deep, unmasked, single-head attention are determined by its input--output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-analytic case of a conjecture of Henry--Marchetti--Kohn, including softmax. We then classify exceptional fibers under explicit normalizer conditions, identifying when collapse makes later scores unobservable, and establish sharp Taylor orders for local identification. Near simultaneous query/key collapse, we compute the complete native Jacobian decay spectrum on separating finite input banks. For common first nonconstant normalizer degree $k$, layer $i$ has contact order $2k3^{i-1}-1$, with exact multiplicities and kernel dimension. High-precision and automatic differentiation calculations illustrate the resulting loss of numerical sensitivity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑