arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少解码器,多编码器:从新视角合成中学习几何表示

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg

arXiv 2610.03717首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SNAP,一种自监督编码器-解码器Transformer,通过姿态条件局部解码器和潜在空间重建目标,从新视角合成中学习几何表示,在多个任务上表现出竞争力。

AI 中文摘要

本文探讨了新视角合成(NVS)在几何表示学习中的作用。原则上,NVS应该推理3D场景结构,从而能够学习可迁移的多视角几何表示。然而,现有的基于编码器的NVS方法产生的表示质量较差。这并非因为缺乏监督信号,而是由于不显眼的架构选择:空间表达性强的解码器稀释了场景编码器的表示能力,以及低层次的像素空间目标阻碍了特征学习。我们提出了SNAP,一种自监督的编码器-解码器Transformer,通过姿态条件局部解码器和潜在空间重建目标来解决这两个问题。SNAP是任务无关的,我们证明它与专门用于几何监督的方法具有竞争力。SNAP还在五个任务上与自监督表示方法竞争:视觉定位、姿态估计、点对应、深度估计和机器人操作。值得注意的是,SNAP的块特征展现出涌现的视角不变性,尽管计算和数据预算较低,却接近重度监督模型。在标准2D表示失效的相机移位下,SNAP的退化更为平缓,表明限制解码器的表达能力有效防止了可迁移几何结构的抑制。此https URL。

英文摘要

This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitive with special-purpose geometry-supervised methods. SNAP also performs competitively against self-supervised representations across five tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch features exhibit emergent viewpoint invariance that approaches heavily supervised models despite lower compute and data budgets. Under camera shifts where standard 2D representations collapse, SNAP degrades more gracefully, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure. https://snap-nvs.github.io

CommentsAccepted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑