arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30592cs.LGcs.CV

QSV:四元数-球面-视觉,用于球面格上的耦合四元数注意力

QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices

Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez

首次发表
浏览论文内容

中文总结 AI 辅助

QSV用单位四元数耦合注意力与特征传输,在球面格上稀疏kNN图传递消息,消融显示传输主导学习,但几何本身不如标准注意力。

中文摘要 AI 辅助

在标准注意力机制中,三个分别学习的投影决定了一个标记对每个邻居的注意力强度($W_Q$、$W_K$)以及被注意特征在聚合前如何变换($W_V$)。我们研究了Quat-Sphere-Vision(QSV),一种稀疏球面视觉模型,该模型用每个标记的单一学习单位四元数取代了这三个投影:相对四元数 $r_{ij} = q_i^{*} \otimes q_j$ 同时提供注意力对数 $\operatorname{Re}(r_{ij})$ 和夹心积特征传输 $x \mapsto r_{ij} \otimes x \otimes r_{ij}^{*}$,消息在同心斐波那契球面上的稀疏kNN图上传递。仅改变目标组件的消融实验显示这两个作用是不对称的。移除特征传输会使CIFAR-10和CIFAR-100上的测试准确率降低约四个百分点(每个CIFAR-100变体单次运行),而将学习的注意力权重替换为均匀平均则基本保持不变。参数匹配的对照实验随后移除了几何本身:同一图上的标准注意力超过了QSV(平均 $87.3\\%$ 对比 $85.9\\%$),而同一模型在平面2D格上达到 $91.1\\%$,与在同一流程下训练的ResNet-20(单次运行)相差 $2.1$ 个百分点。在耦合核中,几乎所有的学习成对计算都存在于特征传输通道中。

英文摘要

In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion $r_{ij} = q_i^{*} \otimes q_j$ supplies both the attention logit $\operatorname{Re}(r_{ij})$ and a sandwich-product feature transport $x \mapsto r_{ij} \otimes x \otimes r_{ij}^{*}$, with messages passed over sparse kNN graphs on concentric Fibonacci spheres. Ablations that change only the targeted component show the two roles to be asymmetric. Removing the transport reduces test accuracy by about four percentage points on CIFAR-10 and CIFAR-100 (single runs per CIFAR-100 variant), while replacing the learned attention weights with uniform averaging leaves it essentially unchanged. Parameter-matched controls then remove the geometry itself: standard attention on the same graph exceeds QSV (mean $87.3\%$ vs. $85.9\%$), and the same model on a flat 2D lattice reaches $91.1\%$, within $2.1$ points of a ResNet-20 trained under the same pipeline (single run). In the coupled kernel, nearly all of the learned pairwise computation resides in the transport channel.

发表机构

  • University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑