arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15018cs.CV

G-ray:相机异构下多视角视觉Transformer中的射线级相对几何位置编码

G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity

Shuo Zhang, Xin Su, Wei Wang, Jun Liu, Xinrui Zeng, Yongsen Chen, Chenjie Wang, Guibo Zhu, Jinqiao Wang, Bin Luo, Liangpei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

G-ray提出射线级相对位置编码,以相机局部射线角度参数化旋转相位,实现投影不变性,在异构相机三维重建中显著降低误差,并提升新视角合成性能。

中文摘要 AI 辅助

我们研究了相机异构条件下多视角视觉Transformer的相对位置编码问题,其中异构性包括视场角(FoV)或投影模型的变化。现有的旋转式相对位置编码通常使用图像平面上的位置坐标,这会产生依赖于投影的相对相位,并为跨投影注意力提供不一致的几何线索。我们引入了G-ray,一种射线级的相对位置编码,其旋转相位由相机局部射线角度参数化。相同的相机局部射线对在不同投影下会产生相同的相对相位,从而提供投影不变的几何一致性。G-ray可以直接使用,也可以与现有编码集成,无需额外学习参数即可保留互补的几何线索。我们在三种宿主编码(RoPE、GTA和RayRoPE)上验证了G-ray,应用于三维重建和新视角合成(NVS)。在50个视图的三个异构三维重建基准上,G-ray在所有六个平均指标上均领先,并将平均点图相对误差比MapAnything降低了45.8%,且两者均提供了校准。仅在均匀针孔图像上训练的3D重建模型,无需重新训练即可处理混合针孔和非针孔输入,并在均匀针孔三维重建协议上保持竞争力。对于NVS,GTA和RayRoPE在联合视角和FoV变化下通过G-ray得到改进。该项目的网页可在以下https URL获取。

英文摘要

We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encodings commonly use image-plane positional coordinates, producing projection-dependent relative phases and inconsistent geometric cues for cross-projection attention. We introduce G-ray, a ray-level relative position encoding whose rotary phases are parameterized by camera-local ray angles. The same camera-local ray pair induces the same relative phase across projections, providing projection-invariant positional consistency. G-ray can be used directly or integrated with existing encodings, retaining complementary geometric cues without additional learned parameters. We validate G-ray in three host encodings, RoPE, GTA, and RayRoPE, across 3D reconstruction and novel-view synthesis (NVS). Across three heterogeneous 3D reconstruction benchmarks at 50 views, G-ray leads all six averaged metrics and reduces mean pointmap relative error by 45.8% over MapAnything, with calibration supplied to both. Trained exclusively on homogeneous pinhole images, the 3D reconstruction model handles mixed pinhole and non-pinhole inputs without retraining and remains competitive on homogeneous pinhole 3D reconstruction protocols. For NVS, GTA and RayRoPE improve with G-ray under joint viewpoint and FoV variation. The project's webpage is available at https://g-ray-project.github.io/.

发表机构

  • Wuhan University(武汉大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • Wuhan AI Research(武汉人工智能研究院)
  • Rongyun Robot (Guizhou) Co., Ltd.(融云机器人(贵州)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑