arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不存在覆盖所有等变注意力的等变架构

No Equivariant Architecture Covers All Equivariant Attention

Tīkun Ông

arXiv 2608.30417首次发表:更新:

AI 中文总结

该研究刻画了等变多头自注意力的性质,证明固定等变MHSA架构会导致表达能力损失,特定群作用下其等变轨迹含大量不可约分支,单一架构仅覆盖其一。

AI 中文摘要

我们对等变多头自注意力(MHSA)给出了完整刻画:若某MHSA层对对称群G是等变的,则G只能通过置换头簇作用,且QK与OV矩阵需满足与该群作用绑定的等变约束。据此我们证明,任何通过多项式参数化无约束MHSA参数实现精确等变的固定MHSA架构,必然会在等变映射类中导致表达能力损失:无约束MHSA的等变轨迹在约化参数空间中形成极多Zariski不可约分支的并集,而任一单一架构最多仅能覆盖其中一个。对于G=D₄作用于C个正则表示副本构成的令牌特征空间,我们表明当注意力头数为8时,存在Ω(C⁶⁴)个分支。

英文摘要

We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$ can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance by polynomially parameterizing unconstrained MHSA parameters inevitably leads to expressivity loss within the class of equivariant maps: the equivariance locus of unconstrained MHSA forms a union of extremely many Zariski-irreducible components in a reduced parameter space, and any single architecture covers at most one. For $G=D_4$ acting on $C$ copies of the regular representation as the token feature space, we show that there are $Ω(C^{64})$ components for eight attention heads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑