发表机构
ETH Zurich; University of Oxford; Norwegian Institute of Bioeconomy Research (NIBIO); University of Zurich(苏黎世联邦理工学院; 牛津大学; 挪威生物经济研究所; 苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探索森林点云基础模型,以LitePT为骨干,通过自监督预训练提升跨任务可迁移性,发现实例判别是主要障碍。
AI 中文摘要
森林清查日益依赖人工智能(AI)模型从大规模三维点云中提取森林属性。当前模型通常专门针对单一任务、传感器和森林类型,这使得在标注、计算和专业知识方面的适应成本高昂。我们探讨是否一个单一的预训练模型能够学习跨不同森林清查设置的可迁移表示。受语言建模和计算机视觉领域最新进展的启发,我们向三维林业的基础模型(FM)迈出了一步。以LitePT为骨干网络,我们首先建立了一个强大的监督基线,在森林语义分割、实例分割、树种分类和年龄回归基准上达到了新的最先进水平。随后,我们整理了一个大规模无标注语料库,涵盖不同森林生态系统中的机载、无人机和移动激光扫描数据,并使用自监督学习对同一骨干网络进行预训练。我们通过比较从零训练、监督预训练和自监督预训练在四个代表性林业任务上、不同标注预算下的表现,系统评估了表示学习策略。与从零训练相比,自监督预训练加速了模型收敛,并在标注稀缺时持续提升性能。与任务特定的监督预训练相比,自监督预训练在下游林业任务中产生了更具可迁移性的表示。这些发现确定了预训练表示最有价值的实际应用场景,并表明实例判别(而非森林语义)是迈向通用三维森林基础模型的主要剩余障碍。代码和模型可在以下网址获取:此https URL。
英文摘要
Forest inventories increasingly rely on artificial intelligence (AI) models to derive forest attributes from large-scale 3D point clouds. Current models are typically specialized to a single task, sensor, and forest type, making adaptation expensive in terms of annotations, computation, and expertise. We ask whether a single pretrained model can instead learn transferable representations across diverse forest inventory settings. Inspired by recent developments in language modelling and computer vision, we take a step toward a foundation model (FM) for 3D forestry. Using LitePT as backbone, we first establish a strong supervised baseline that sets a new state of the art on forest semantic and instance segmentation, tree species classification, and age regression benchmarks. We then curate a large-scale unlabelled corpus spanning airborne, UAV, and mobile laser scanning across diverse forest ecosystems, and pretrain the same backbone using self-supervised learning. We systematically evaluate representation learning strategies by comparing training from scratch, supervised pretraining, and self-supervised pretraining across four representative forestry tasks, under varying annotation budgets. Compared with training from scratch, self-supervised pretraining accelerates model convergence and consistently improves performance when annotations are scarce. Compared with task-specific supervised pretraining, self-supervised pretraining yields more transferable representations across downstream forestry tasks. These findings identify the practical regime in which pretrained representations are most valuable and suggest that instance discrimination, rather than forest semantics, is the main remaining obstacle to a general-purpose 3D forest foundation model. Code and models are available at: https://github.com/prs-eth/ForPT.
CommentsProject page: https://prs-eth.github.io/ForPT