发表机构
Technical University of Munich; Siemens AG; Munich Center for Machine Learning; ROBOX(慕尼黑工业大学; 西门子股份公司; 慕尼黑机器学习中心; ROBOX)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LEGAU提出统一框架,利用语义高斯先验联合预测NOCS对应、姿态和尺寸,通过多模态融合提升类别级6D姿态估计,在SOPE上提升22%。
AI 中文摘要
从单次RGB-D观测中进行类别级6D姿态估计本质上是欠约束的,因为部分可见的几何结构必须与规范对象结构一起解释,才能确定稳定的姿态。我们提出了LEGAU,一个统一框架,联合预测NOCS对应关系、对象姿态和尺寸,以及一个规范语义高斯场。LEGAU不将重建视为独立的辅助任务,而是将高斯场作为类别条件结构先验,参与多模态特征融合,并为局部姿态推理提供全局指导。在类别文本嵌入的条件下,LEGAU通过基于Transformer的融合模块处理RGB-D观测,该模块整合视觉、几何和类别级线索,解码NOCS图、姿态和尺寸信息以及基于高斯的目标表示。在合成和真实世界基准上的大量实验表明,这种耦合的姿态-形状公式在单模型多类别设置中实现了强性能,在SOPE上提升高达22%,并具有竞争力的真实世界数据迁移能力。这些结果凸显了在统一表示中联合学习规范对应、对象形状和姿态对齐的优势。
英文摘要
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object pose and size, and a canonical Semantic Gaussian Field. Rather than treating reconstruction as a detached auxiliary task, LEGAU uses the Gaussian field as a category-conditioned structural prior that participates in multimodal feature fusion and provides global guidance for local pose reasoning. Conditioned on a categorical text embedding, LEGAU processes RGB-D observations through a transformer-based fusion module that integrates visual, geometric, and category-level cues, decoding the NOCS map, pose and size information and the Gaussian-based object representation. Extensive experiments on synthetic and real-world benchmarks show that this coupled pose-shape formulation achieves strong performance in a single-model multi-category setting, with up to 22\% on SOPE and competitive transfer to real-world data. These results highlight the benefit of jointly learning canonical correspondence, object shape, and pose alignment within a unified representation.