arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01589cs.CV

PAGER:通过几何与关系蒸馏实现的部分到全局对齐

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni

首次发表
浏览论文内容

中文总结 AI 辅助

针对部分观察与全局3D特征空间不匹配问题,提出无标签适应方法PAGER,通过几何与关系蒸馏对齐特征,在零样本迁移中超越完全微调模型。

中文摘要 AI 辅助

预训练的3D编码器通常在全局重建的场景上开发,这些场景以一致的世界坐标框架表示,而具身系统必须从相机坐标中部分、依赖于视点的观察进行推理。我们表明,从全局学习的3D特征空间到现实的部分观察的转变暴露了严重的表示不匹配,我们在代表性的最先进编码器中一致地发现了这一点,包括Sonata和Concerto。冻结的Sonata编码器在完整ScanNet场景上使用全局线性探针实现了72.47 mIoU,但在单帧相机坐标输入上仅达到2.57 mIoU。无需训练的引力对齐将性能恢复到41.64 mIoU,表明坐标框架不匹配是性能下降的主要来源,但仅通过规范化无法完全解决。我们引入了PAGER,一种无标签适应方法,仅使用配对的部分/全局几何将部分视图特征与冻结的全局3D语义空间对齐。它学习轻量级适应模块,同时保持预训练编码器和全局分割探针冻结。匹配点特征对齐将部分特征锚定到其全局对应物,而关系监督保留了它们相对于全局表示的相似性结构。全局几何仅在训练期间提供监督。推理直接对部分观察进行操作。在没有部分视图标签的情况下,PAGER在Sonata和Concerto上均优于标签监督的PEFT,并且在零样本ScanNet到ScanNet++迁移中超过了完全微调的Sonata(53.93 vs. 48.09 mIoU),表明保留冻结的全局表示可以改善跨数据集迁移。

英文摘要

Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial observations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label-free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmentation probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides supervision only during training. Inference operates directly on the partial observation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet$\rightarrow$ScanNet++ transfer surpasses fully fine-tuned Sonata ($53.93$ vs.\ $48.09$ mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer.

发表机构

  • Technical University of Denmark(丹麦技术大学)
  • Pioneer Center for Artificial Intelligence(先锋人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

↑