arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

矩阵 zonotopic 注意力:面向集合 Transformer 的上下文自适应值投影

Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers

Zhen Zhang, Amr Alanwar

arXiv 2608.05472首次发表:更新:

AI 中文总结

该研究针对集合 Transformer 提出矩阵 zonotopic 注意力(MZAttn),通过上下文自适应值投影解决多头注意力的不对称性问题,在高秩稀疏组合的集合预测任务上表现出架构优势。

AI 中文摘要

多头注意力将依赖输入的 softmax 路由与不依赖输入的线性值投影相结合,因此每个输入集合对应的将聚合值映射到输出的逐样本算子是相同的。我们研究了这种不对称性对置换不变集合目标的影响。我们引入目标算子的变换自由度(TDOF),这是一种衡量精确表示所需的依赖输入方向数量的复杂度指标,并提出深度分离分析表明,上下文刚性注意力所需的深度与目标的 TDOF 成正比,而具有上下文自适应值族的单层即可表示相同目标。基于此分析,我们提出矩阵 zonotopic 注意力(MZAttn),它用上下文自适应矩阵 zonotope 族替代固定值投影:由一个中心矩阵加上生成矩阵的加权和(权重为依赖输入的门控)构成。该构造在初始化时退化为标准多头注意力,保持置换等变性,并具有数据驱动的可达性解释。在一系列集合预测任务上的实验与 TDOF 预测一致,即架构优势是有选择性的:当目标以高秩、稀疏组合的方式依赖输入集合时,该优势显现;而在参数匹配的标准注意力已具竞争力的聚合统计目标上,优势很小。

英文摘要

Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set. We study the consequences of this asymmetry for permutation-invariant set targets. We introduce the Transformation Degrees of Freedom (TDOF) of a target operator, a complexity measure counting the input-dependent directions an exact representation requires, and present a depth-separation analysis showing that context-rigid attention needs depth proportional to the target's TDOF, whereas a single layer with a context-adaptive value family can represent the same target. Building on this analysis, we propose Matrix Zonotopic Attention (MZAttn), which replaces the fixed value projection with a context-adaptive matrix-zonotope family: a centre matrix plus a sum of generator matrices weighted by input-dependent gates. The construction reduces to standard multi-head attention at initialisation, preserves permutation equivariance, and admits a data-driven reachability interpretation. Experiments on a range of set-prediction tasks are consistent with the TDOF prediction that the architectural advantage is selective: it appears on targets that depend on the input set in a high-rank, sparsely combinatorial way, and is small on aggregate-statistic targets where parameter-matched standard attention is already competitive.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑