arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先精炼后融合:基于优先级精炼与多模态知识融合的无训练三维点云适配

Refine Then Fusion: Training-Free 3D Point Cloud Adaptation with Priority Refinement and Multi-Modal Knowledge Fusion

Hang Cheng, Yan Chen, Mingyu Fan, Long Zeng

arXiv 2609.21522首次发表:更新:

发表机构

Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对少样本三维点云识别中多模态特征冗余与模态可靠性差异问题,提出无训练框架RTF,通过优先级通道精炼和可靠性感知融合实现自适应多模态聚合,在五个基准上达到最先进性能。

AI 中文摘要

近期预训练基础模型为下游三维视觉任务提供了丰富的多模态先验知识。然而,这些表示在少样本场景中的有效性受到两个基本挑战的限制:高维特征通常包含大量通道冗余和与任务无关的噪声,而不同模态的可靠性在不同样本间存在差异。因此,直接聚合异构表示会忽略样本相关的模态可靠性,并可能掩盖关键的判别性线索。为解决这些局限,我们提出“先精炼后融合”(Refine Then Fusion, RTF),一种用于少样本三维识别的无训练框架。RTF首先通过联合建模类间相似性和类内稳定性来识别判别性特征通道,从而将领域特定知识的精炼与预训练模型的缓存表示解耦。随后,它引入一种可靠性感知的融合机制,通过特征精炼引起的分布偏移来估计样本级模态可靠性,从而实现多模态表示的自适应聚合。此外,RTF构建了一个记忆缓存,将实例级支持特征与类级原型相结合以推断查询标签。在五个基准上的大量实验表明,RTF始终优于单模态基线、部分融合变体和现有轻量级适配方法,在无需梯度优化、额外训练数据、辅助训练或参数更新的情况下,实现了最先进的少样本三维识别性能。

英文摘要

Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representations in few-shot scenarios is limited by two fundamental challenges: High-dimensional features often contain substantial channel redundancy and task-irrelevant noise, while the reliability of different modalities varies across samples. Consequently, direct aggregation of heterogeneous representations overlooks sample-dependent modality reliability and may obscure the discriminative cues essential. To address these limitations, we propose Refine Then Fusion(RTF), a training-free framework for few-shot 3D recognition. RTF first identifies discriminative feature channels by jointly modeling inter-class similarity and intra-class stability, thereby decoupling domain-specific knowledge refinement from the cached representations of pre-trained models. It then introduces a reliability-aware fusion mechanism that estimates sample-wise modality reliability from the distribution shifts induced by feature refinement, enabling adaptive aggregation of multi-modal representations. Furthermore, RTF constructs a memory cache that integrates instance-level support features with class-level prototypes to infer query labels. Extensive experiments on five benchmarks demonstrate that RTF consistently outperforms single-modal baselines, partial-fusion variants, and existing lightweight adaptation methods, achieving state-of-the-art few-shot 3D recognition performance without gradient optimization, additional training data, auxiliary training, or parameter updates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑