发表机构
Bowdoin College(鲍登学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有多模态3D预训练方法需多模型适配不同计算预算的问题,提出3D-MRL框架,可通过单个模型生成不同维度3D表示,在多数据集上提升了3D识别与检索性能。
AI 中文摘要
视觉-语言模型将点云与图像和文本嵌入对齐,实现了3D形状的零样本识别、检索和开放词汇理解。现有多模态3D预训练方法生成固定维度的嵌入,需为不同计算预算使用单独模型。我们提出3D套娃表示学习框架(3D-MRL),该框架基于套娃表示学习,通过将点云与冻结的CLIP图像和文本嵌入对齐,同时在多个嵌入维度上应用对比监督,学习嵌套的3D表示。套娃目标仅应用于3D编码器,使单个模型无需重新训练即可生成不同维度的表示。在Objaverse-LVIS、ModelNet40和ScanNet数据集上的实验表明,3D-MRL在零样本和少样本3D识别任务中取得了有竞争力的性能,此外,学习到的表示支持单个模型内不同嵌入维度间的检索。在Objaverse-LVIS上,3D-MRL将Top-1准确率从46.8%提升至50.9%,检索实验进一步显示,不同嵌入维度产生不同水平的语义和几何特异性。
英文摘要
Vision-Language Models align point clouds with image and text embeddings, enabling zero-shot recognition, retrieval, and open-vocabulary understanding of 3D shapes. Existing multimodal 3D pre-training methods produce fixed-dimensional embeddings, requiring separate models for different computational budgets. We propose 3D Matryoshka Representation Learning (3D-MRL), a multimodal 3D pre-training framework based on Matryoshka Representation Learning. 3D-MRL learns nested 3D representations by aligning point clouds with frozen CLIP image and text embeddings while applying contrastive supervision across multiple embedding dimensions. The Matryoshka objective is applied only to the 3D encoder, allowing a single model to produce representations at different dimensionalities without retraining. Experiments on the Objaverse-LVIS, ModelNet40, and ScanNet datasets show that 3D-MRL achieves competitive performance on zero-shot and few-shot 3D recognition tasks. In addition, the learned representations support retrieval across different embedding dimensions within a single model. On Objaverse-LVIS, 3D-MRL improves Top-1 accuracy from 46.8% to 50.9%. Retrieval experiments further show that different embedding dimensions yield varying levels of semantic and geometric specificity.
CommentsAccepted at the British Machine Vision Conference (BMVC) 2026. Project page: https://usmarcv.github.io/3DMRL_page/