发表机构
Durham University; University of Bristol; Tsinghua University; University of Chinese Academy of Sciences(杜伦大学; 布里斯托大学; 清华大学; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3D视觉-语言任务中异构表示融合困难的问题,提出先三重对齐后正交约束融合的框架,在八个数据集上显著提升分割、定位等性能。
AI 中文摘要
统一的3D视觉-语言系统必须结合互补的几何、尺度和外观线索,同时支持从实例分割到语言引导推理的任务。现有方法通常独立处理点云、体素网格和多视角图像;直接组合这些异构表示可能会留下大量特征差异未解决,而后续的无约束适应可能会扭曲其内部几何结构。我们提出了一种先对齐后融合的框架,该框架首先应用三重成对余弦对齐,以在三种表示之间建立片段级对应关系,然后通过提示引导的查询解码器检索任务条件特征。在融合之前,特定表示的查询特征通过约束在特殊正交群上的可学习映射进行变换。这些映射保留了每种表示内的内积和欧几里得距离,允许受控的表示特定重新参数化,而不会任意扭曲其内部几何结构。变换后的特征随后在下游任务监督下通过自适应融合进行组合。实验涵盖了八个数据集,用于实例分割、视觉定位、问答和密集描述。与PQ3D相比,该模型在ScanNet200上的平均精度提高了3.2个百分点,在ScanRefer、Nr3D、Sr3D和Multi3DRefer上的定位准确率分别提高了2.9、10.6、4.6和4.1个百分点,同时在ScanQA、SQA3D和Scan2Cap上也提高了性能。消融研究进一步支持了对齐和正交重新参数化的互补作用以及自适应融合的有效性。
英文摘要
Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous representations may leave substantial feature discrepancy unresolved, while subsequent unconstrained adaptation may distort their internal geometry. We propose an align-then-fuse framework that first applies triple pairwise cosine alignment to establish segment-level correspondence across the three representations and then retrieves task-conditioned features with a prompt-guided query decoder. Before fusion, representation-specific query features are transformed by learnable mappings constrained to the special orthogonal group. These mappings preserve inner products and Euclidean distances within each representation, permitting controlled representation-specific re-parameterisation without arbitrarily distorting its internal geometry. The transformed features are subsequently combined through Adaptive Fusion under downstream task supervision. Experiments cover eight datasets for instance segmentation, visual grounding, question answering, and dense captioning. Compared with PQ3D, the model improves average precision by 3.2 points on ScanNet200 and grounding accuracy by 2.9, 10.6, 4.6, and 4.1 points on ScanRefer, Nr3D, Sr3D, and Multi3DRefer, respectively, while also improving performance on ScanQA, SQA3D, and Scan2Cap. Ablations further support the complementary roles of alignment and orthogonal re-parameterisation and the effectiveness of Adaptive Fusion.