arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32899cs.CV

重新思考极坐标下的全双曲视觉Transformer

Rethinking the Fully Hyperbolic Vision Transformer in Polar Coordinates

Ahmad Bdeir, Niels Landwehr

首次发表
浏览论文内容

中文总结 AI 辅助

针对双曲变换器在环境坐标下数值不稳定、限制半径利用的问题,提出极坐标下的全双曲变换器,含极坐标全连接层、等距双曲移位和平均半径Lorentz提升,在视觉任务上显著优于基线,并能在ImageNet上预训练后泛化。

中文摘要 AI 辅助

双曲空间凭借其体积随距原点距离呈指数增长的特性,能够以低失真嵌入数据中的内在层级结构。然而,当前的Lorentz变换器模块是在环境坐标中构建的,在较大半径处,由于Lorentz内积的不稳定性,数值误差会增大,导致许多模型限制半径以避免此问题。这阻碍了我们利用双曲空间中那些激励该几何结构的区域。为解决此问题,我们重新审视了极坐标下的变换器模块组件,其中诸如距离计算和注意力质心等双曲运算可以在环境坐标公式中无数值抵消的情况下进行计算。具体而言,我们提出了一种极坐标全连接层,该层分别映射嵌入的方向和半径,使得特征的半径可以被学习,而非由线性映射的范数决定。我们还引入了等距双曲移位作为相对位置编码,其查询-键距离随令牌间隔呈对数增长。最后,我们将残差连接重构为平均半径Lorentz提升。结合这些组件,我们开发了一种全双曲变换器,在标准视觉任务上显著优于欧几里得和双曲基线。我们进一步在ImageNet上评估了我们的模型,并展示了其利用预训练权重泛化到其他数据集的能力,这与欧几里得对应模型类似。

英文摘要

Hyperbolic space can embed intrinsic hierarchies in data with low distortion due to the exponential growth of volume with distance from the origin. However, current Lorentz transformer blocks are formulated in ambient coordinates, where numerical errors increase at large radii due to instability in the Lorentzian inner product, leading many models to limit the radius to avoid this issue. This prevents us from utilizing the regions of hyperbolic space that motivate the geometry. To address this, we revisit the components of the transformer block in polar coordinates, where hyperbolic operations such as distance calculation and attention centroids can be computed without the numerical cancellation in their ambient-coordinate formulations. Specifically, we propose a polar fully connected layer that separately maps an embedding's direction and radius, allowing the radius of a feature to be learned rather than determined by the norm of a linear map. We additionally introduce horospherical shifts as relative positional encodings whose query-key distances grow logarithmically with the token gap. Finally, we reformulate the residual connection as average radius Lorentz boosts. Combining these components, we develop a fully hyperbolic transformer that substantially improves performance over Euclidean and hyperbolic baselines on standard vision tasks. We further evaluate our model on ImageNet and demonstrate its ability to generalize to other datasets using pre-trained weights, similar to Euclidean counterparts.

↑