用于旋转稳定的360°场景理解的规范等变注意力
Gauge-Equivariant Attention for Rotation-Stable $360^\circ$ Scene Understanding
浏览论文内容
中文总结 AI 辅助
提出规范等变相对位置编码(GE-RPE)解决球形模型旋转不稳定性,结合教师令牌置换和iBOT+MAE预训练,实现360°场景理解的高精度低标签微调。
中文摘要 AI 辅助
全景360°场景理解日益依赖于icosphere变换器,但最先进的球形模型在相机旋转90°时,其分割精度会损失超过一半,且受控消融实验表明,其相对位置偏置中的规范依赖性是一个主要因素。我们提出了规范等变相对位置编码(GE-RPE):一种对有限循环子群C_n⊂SO(2)的规范旋转上的偏置进行无参数Reynolds平均的方法。将其插入SphereUFormer骨干网络后,部署时变化不可见——零额外参数,前向延迟增加1.5%–4.4%——且匹配的三种子GE-RPE模型记录了1.3%的下降;已发布的SphereUFormer检查点在相同的压力协议下记录了53%的下降,但使用了不同的训练方案。一旦消除规范缺陷,并通过教师令牌置换π_R将SSL视图对齐到旋转的学生框架,iBOT+MAE预训练就不再是负担,而成为干净的低标签杠杆:完整框架EquiSSL(GE-RPE + π_R + iBOT+MAE)将下降收紧至0.8%,在68.30%的验证mIoU下,并在N=373的测试分割上将1%标签微调提升了+2.39 mIoU(在较小的N=40验证分割上提升了+4.10);相同的修复也适用于单目深度估计和零样本Structured3D分割。该构造可证明是C_n不变的,并且与连续的SO(2)平均值的差距为O(n^{-2}),使得所得模型成为可用的360°视觉计算原语,适用于全景重光照、沉浸式视频和跨数据集迁移。代码可在https URL获取。
英文摘要
Panoramic $360^\circ$ scene understanding increasingly relies on icosphere transformers, but a state-of-the-art spherical model loses more than half of its segmentation accuracy when the camera rotates by $90^\circ$, and controlled ablations identify gauge dependence in its relative-position bias as a major contributor. We propose gauge-equivariant relative position encoding (GE-RPE): a parameter-free Reynolds average of the bias over a finite cyclic subgroup $C_n\!\subset\!\mathrm{SO}(2)$ of gauge rotations. Plugged into a SphereUFormer backbone the change is invisible at deployment---zero added parameters and $1.5$--$4.4\%$ forward latency---and the matched three-seed GE-RPE model records a $1.3\%$ drop; the published SphereUFormer checkpoint records $53\%$ under the same stress protocol but a different training recipe. Once the gauge defect is removed and a teacher-token permutation $π_R$ aligns the SSL views to the rotated student frame, iBOT$+$MAE pretraining stops being a liability and becomes a clean low-label lever: the full framework EquiSSL (GE-RPE $+$ $π_R$ $+$ iBOT$+$MAE) tightens the drop to $0.8\%$ at $68.30\%$ val mIoU and lifts $1\%$-label fine-tuning by $+2.39$ mIoU on the $N{=}373$ test split (and by $+4.10$ on the smaller $N{=}40$ val split); the same fix carries over to monocular depth and to zero-shot Structured3D segmentation. The construction is provably $C_n$-invariant and $\mathcal{O}(n^{-2})$-close to the continuous $\mathrm{SO}(2)$ average, making the resulting model a usable $360^\circ$ visual-computing primitive across panoramic relighting, immersive video, and cross-dataset transfer. Code is available at https://github.com/Jaywalk18/equissl-release.