arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TOLiD:通过用于蒸馏的令牌提升弥合视觉基础模型与激光雷达预训练之间的架构差距

TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation

Sutharsan Mahendran, Darshana Priyasad, Kaushik Roy, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam

arXiv 2607.10762首次发表:更新:

发表机构

SAIVT Group, School of Electrical Engineering and Robotics, Queensland University of Technology; CSIRO Robotics, CSIRO(澳大利亚昆士兰科技大学电气工程与机器人学院SAIVT组; 澳大利亚联邦科学与工业研究组织机器人部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出TOLiD方法,通过耦合激光雷达主干与从冻结VFM教师初始化的学生ViT,利用视锥体池化、注意力及掩码双线性采样等技术,解决模态和架构差距问题,在多数据集上评估显示能提升迁移效果。

AI 中文摘要

从视觉基础模型(VFM)到激光雷达主干的跨模态蒸馏最近成为一种自监督预训练策略,可减少对3D场景理解的密集逐点注释的依赖。然而,现有蒸馏管道通常将VFM视为固定特征源,并训练异构3D主干以匹配固定图像嵌入,迫使学生弥合密集ViT令牌表示与稀疏3D编码器之间的模态差距和跨架构差距。我们提出了TOLiD,一种用于激光雷达表示学习的自监督预训练方法,通过将激光雷达主干与从冻结的VFM教师初始化的学生视觉Transformer(ViT)耦合,并对兼容的补丁令牌表示应用监督来解决这一差距。TOLiD使用视锥体池化和视锥体注意力将每个图像补丁视锥体内的点特征集转换为一个令牌,并通过可见性掩码执行令牌级蒸馏。对于仅激光雷达部署,我们使用掩码双线性采样将令牌特征提升回逐点表示,以避免激光雷达点有限的补丁。我们在五个异构激光雷达数据集和四个跨传感器适应对上广泛评估了TOLiD,展示了在冻结主干和轻量级头部情况下改进的迁移效果。

英文摘要

Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged as a self-supervised pretraining strategy that reduces reliance on dense point-wise annotation for 3D scene understanding. However, existing distillation pipelines typically treat the VFM as a frozen feature source and train a heterogeneous 3D backbone to match fixed image embeddings, forcing the student to bridge both the modality gap and the cross-architecture gap between dense ViT token representations and sparse 3D encoders. We propose TOLiD, a self-supervised pretraining method for LiDAR representation learning that addresses this gap by coupling a LiDAR backbone with a student Vision Transformer (ViT) initialized from a frozen VFM teacher and applying supervision over compatible patch-token representations. TOLiD converts the set of point features within each image patch frustum into a token using Frustum Pooling followed by Frustum Attention, and performs token-level distillation with visibility masking. For LiDAR-only deployment, we lift token features back to per-point representations using masked bilinear sampling to avoid patches that have limited LiDAR points. We extensively evaluate TOLiD on five heterogeneous LiDAR datasets and four cross-sensor adaptation pairs, demonstrating improved transfer with frozen backbones and lightweight heads.

CommentsAccepted to The IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑