用于6D位姿估计中条件流匹配的基础特征融合
Foundational feature fusion for conditional flow matching in 6D pose estimation
浏览论文内容
中文总结 AI 辅助
本研究针对6D位姿估计中现有方法需训练任务特定编码器、融合策略简单的问题,提出FunFlow6D,利用几何与外观基础模型特征及交叉注意力融合机制,在BOP四个数据集上实现最优性能,降低了监督需求和内存开销。
中文摘要 AI 辅助
条件流匹配已推动物体6D位姿估计取得进展,通过逐步去噪并将物体表示与观测场景配准,实现了当前最优性能。现有方法需要在物体-场景重叠监督下训练任务特定编码器,且依赖简单的特征融合策略来解决位姿歧义问题。本文提出FunFlow6D,一种新型基于流匹配的公式,利用几何和外观基础模型的特征进行位姿估计,无需任务特定编码器训练;还引入基于交叉注意力的融合机制,动态结合几何与外观特征,为流匹配模块提供更丰富的条件。在BOP基准的四个数据集上实验表明,FunFlow6D在超越先前最优性能的同时,降低了监督需求和内存开销,大量 ablation 验证了各提出组件的贡献。项目网站:this https URL。
英文摘要
Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task-specific encoders supervised on object-scene overlap and rely on trivial feature fusion strategies to resolve pose ambiguities. We present FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoder training. We also introduce a cross attention-based fusion mechanism that dynamically combines geometric and appearance features to provide richer conditioning for the flow matching module. Experiments on four datasets from the BOP benchmark show that FunFlow6D outperforms the previous state of the art while reducing supervision requirements and memory overhead. Extensive ablations validate the contribution of each proposed component. Project website: https://tev-fbk.github.io/FunFlow6D/.
发表机构
- Fondazione Bruno Kessler(布鲁诺·凯塞勒基金会)
- University of Trento(特伦托大学)
机构由 AI 辅助整理,请以论文原文为准。