发表机构
University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出MeshFM框架,将2D视觉基础模型的特征提炼到3D,无需3D标注和推理优化,其2D监督训练的特征可直接用于多项3D下游任务,性能媲美3D监督方法且对极端旋转鲁棒。
AI 中文摘要
我们提出了MeshFM,一种用于从3D输入中提取丰富特征的高效前馈框架。我们的方法将视觉基础模型中的2D特征提炼到3D空间,训练一个前馈网络直接预测3D特征,推理过程无需优化。该方法采用两阶段训练策略:首先仅使用2D特征监督优化3D中的特征场;其次训练网络回归该特征场。整个过程无需3D标注,仅依赖2D基础模型中的强大信息。我们证明,所学习的特征可直接应用于下游任务,包括部件分割、密集对应和网格变形。大量实验表明,仅用2D监督训练的MeshFM,性能与明确使用3D监督训练的方法相当,甚至无需特定任务微调,且模型对输入对象的极端旋转具有鲁棒性。项目页面:this https URL
英文摘要
We present MeshFM, an efficient feedforward framework for extracting rich features from 3D inputs. Our method distills 2D features from visual foundation models into 3D. We train a feedforward network to directly predict 3D features without requiring optimization during inference. The approach utilizes a two-stage training strategy. First, we optimize a feature field in 3D using only 2D feature supervision. Second, we train a network to regress this feature field. The entire procedure requires no 3D annotation, instead relying on the powerful information in 2D foundation models. We demonstrate that our learned features can be immediately applied to downstream tasks, including part segmentation, dense correspondence, and mesh deformation. Extensive experiments show that MeshFM, trained solely with 2D supervision, performs on par with methods trained explicitly with 3D supervision, even without task-specific fine-tuning. Moreover, our model is trained to be robust to extreme rotations of the input objects. Project page: https://threedle.github.io/MeshFM/
CommentsProject page: https://threedle.github.io/MeshFM/