G-ray:相机异构下多视角视觉Transformer中的射线级相对几何位置编码
G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity
浏览论文内容
中文总结 AI 辅助
G-ray提出射线级相对位置编码,以相机局部射线角度参数化旋转相位,实现投影不变性,在异构相机三维重建中显著降低误差,并提升新视角合成性能。
中文摘要 AI 辅助
我们研究了相机异构条件下多视角视觉Transformer的相对位置编码问题,其中异构性包括视场角(FoV)或投影模型的变化。现有的旋转式相对位置编码通常使用图像平面上的位置坐标,这会产生依赖于投影的相对相位,并为跨投影注意力提供不一致的几何线索。我们引入了G-ray,一种射线级的相对位置编码,其旋转相位由相机局部射线角度参数化。相同的相机局部射线对在不同投影下会产生相同的相对相位,从而提供投影不变的几何一致性。G-ray可以直接使用,也可以与现有编码集成,无需额外学习参数即可保留互补的几何线索。我们在三种宿主编码(RoPE、GTA和RayRoPE)上验证了G-ray,应用于三维重建和新视角合成(NVS)。在50个视图的三个异构三维重建基准上,G-ray在所有六个平均指标上均领先,并将平均点图相对误差比MapAnything降低了45.8%,且两者均提供了校准。仅在均匀针孔图像上训练的3D重建模型,无需重新训练即可处理混合针孔和非针孔输入,并在均匀针孔三维重建协议上保持竞争力。对于NVS,GTA和RayRoPE在联合视角和FoV变化下通过G-ray得到改进。该项目的网页可在以下https URL获取。
英文摘要
We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encodings commonly use image-plane positional coordinates, producing projection-dependent relative phases and inconsistent geometric cues for cross-projection attention. We introduce G-ray, a ray-level relative position encoding whose rotary phases are parameterized by camera-local ray angles. The same camera-local ray pair induces the same relative phase across projections, providing projection-invariant positional consistency. G-ray can be used directly or integrated with existing encodings, retaining complementary geometric cues without additional learned parameters. We validate G-ray in three host encodings, RoPE, GTA, and RayRoPE, across 3D reconstruction and novel-view synthesis (NVS). Across three heterogeneous 3D reconstruction benchmarks at 50 views, G-ray leads all six averaged metrics and reduces mean pointmap relative error by 45.8% over MapAnything, with calibration supplied to both. Trained exclusively on homogeneous pinhole images, the 3D reconstruction model handles mixed pinhole and non-pinhole inputs without retraining and remains competitive on homogeneous pinhole 3D reconstruction protocols. For NVS, GTA and RayRoPE improve with G-ray under joint viewpoint and FoV variation. The project's webpage is available at https://g-ray-project.github.io/.
发表机构
- Wuhan University(武汉大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Wuhan AI Research(武汉人工智能研究院)
- Rongyun Robot (Guizhou) Co., Ltd.(融云机器人(贵州)有限公司)
机构由 AI 辅助整理,请以论文原文为准。