arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向四维点云视频的掩码点管联合嵌入预测自监督学习

Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

Jheng-Ling Lee, Shang-Tse Chen

arXiv 2608.24093首次发表:更新:

发表机构

National Taiwan University(台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对四维点云视频自监督学习的挑战,提出JEPA风格框架结合Sketched Isotropic Gaussian Regularization,通过掩码时空区域的潜在点管预测学习表示,在动作和手势识别任务中取得良好效果。

AI 中文摘要

四维点云视频的自监督表示学习颇具挑战,因为标注成本高昂,而基于重建的预训练可能过度强调低级几何细节。我们提出一种JEPA风格框架,通过潜在点管预测从未标注的时空点云中学习。该模型不重建原始坐标,而是掩码时空区域,并从特征空间中的可见上下文表示预测其目标表示。为稳定潜在预测,我们引入Sketched Isotropic Gaussian Regularization,该方法无需依赖显式重建目标即可鼓励非塌陷嵌入。此公式旨在捕捉空间结构与时间动态,同时保持预训练目标与下游语义识别一致。在动作和手势识别基准上的实验表明,所学表示可提升下游微调、有限标签学习及跨数据集迁移效果。这些结果表明,JEPA风格潜在预测是面向四维点云视频的以重建为中心预训练的有前景替代方案。

英文摘要

Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining can overemphasize low-level geometric details. We propose a JEPA-style framework that learns from unlabeled spatiotemporal point clouds through latent point-tube prediction. Instead of reconstructing raw coordinates, the model masks spatiotemporal regions and predicts their target representations from visible context representations in feature space. To stabilize latent prediction, we incorporate Sketched Isotropic Gaussian Regularization, which encourages non-collapsed embeddings without relying on explicit reconstruction targets. This formulation aims to capture both spatial structure and temporal dynamics while keeping the pretraining objective aligned with downstream semantic recognition. Experiments on action and gesture recognition benchmarks show that the learned representations improve downstream fine-tuning, limited-label learning, and cross-dataset transfer. These results suggest that JEPA-style latent prediction is a promising alternative to reconstruction-centered pretraining for 4D point cloud videos.

Comments13 pages, 4 figures; supplementary material included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑