arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrevasseSeg:一种标签高效的无人机冰裂缝分割框架

CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework

Steven Wallace, William D. Harcourt, Richard Hann, Aiden Durrant, Somayajulu Sripada, Georgios Leontidis

arXiv 2608.15790首次发表:更新:

发表机构

School of Natural and Computing Sciences, University of Aberdeen; Interdisciplinary Institute, University of Aberdeen; School of Geosciences, University of Aberdeen; Department of Engineering Cybernetics, Norwegian University of Science and Technology; School of Computing Sciences, University of East Anglia; Department of Physics and Technology, UiT The Arctic University of Norway(阿伯丁大学自然与计算科学学院; 阿伯丁大学跨学科研究所; 阿伯丁大学地球科学学院; 挪威科技大学工程控制论系; 东英吉利大学计算科学学院; 北极挪威大学物理与技术系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CrevasseSeg框架针对无人机冰裂缝分割,采用多种自监督目标与架构,发现DINOv3特征在不同读出方式下表现反转,其优化 pipeline 性能优于基线,支持遥感标签高效分割研究。

AI 中文摘要

从无人机(UAV)影像绘制冰裂缝图对冰川学研究和冰川地形的野外安全至关重要。然而,冰川表面的像素级标注成本高昂,且需要领域专家。我们提出了CrevasseSeg,这是一个针对斯瓦尔巴群岛博雷布雷恩(Borebreen)冰舌的二值分割框架,包含1938张未标注的无人机正射镶嵌图块用于自监督/无监督微调、24张标注图块用于验证、176张标注图块用于测试。使用CrevasseSeg,我们在三种架构(O-Net、O-Net++和以DINOv3初始化的O-Net)上对五种自监督目标——BYOL、詹森-香农散度(JSD)目标、Barlow-Twins、VICReg以及组合的BYOL-JSD目标——进行了基准测试。每种配置在两种冻结特征读出方式下进行评估,这两种读出方式仅在决策边界形式上不同:仅基于24张标注验证图像拟合的线性探测和非线性XGBoost分类器。我们的核心发现是两种读出方式之间存在一致的反转:DINOv3特征在线性探测下最弱,但在非线性读出下最强。对学习到的特征空间的UMAP分析显示,DINOv3将像素分割为许多小簇,其中类别局部交错,而卷积架构(O-Net和O-Net++)将像素嵌入到单一的类别排序流形上。基于卫星预训练的DINOv3在所有目标上均优于自然图像初始化,我们的标签高效DINOv3-ViT-L-Sat-O-Net-BYOL-JSD pipeline达到75.33 mDSC / 61.28 mIoU,优于以RGB像素值作为特征、在相同24张标注图像上拟合的标准机器学习基线。我们发布CrevasseSeg以支持遥感中的标签高效分割研究。

英文摘要

Crevasse mapping from uncrewed aerial vehicle (UAV) imagery matters for glaciological research and for field safety in glaciated terrain. Yet, pixel-level annotation of glacier surfaces is costly and requires domain experts. We introduce CrevasseSeg, a framework for binary segmentation over the terminus of Borebreen, Svalbard, comprising 1,938 unlabelled UAV orthomosaic tiles for self-supervised/unsupervised fine-tuning, 24 labelled tiles for validation and 176 labelled tiles for testing. Using CrevasseSeg, we benchmark five self-supervised objectives -- BYOL, a Jensen-Shannon Divergence (JSD) objective, Barlow-Twins, VICReg, and a combined BYOL-JSD objective -- across three architectures: O-Net, O-Net++, and a DINOv3-initialised O-Net. Each configuration is evaluated under two frozen-feature readouts that differ only in the form of their decision boundary: a linear probe and a non-linear XGBoost classifier fit only on the 24 labelled validation images. Our central finding is a consistent inversion between the two readouts: DINOv3 features are the weakest under linear probing but the strongest under a non-linear readout. A UMAP analysis of the learned feature space shows that DINOv3 fragments pixels into many small clusters in which the classes are locally interleaved, whereas the convolutional architectures (O-Net and O-Net++) embed them onto a single class-sorted manifold. Satellite-pretrained DINOv3 improves over natural-image initialisation across objectives, and our label-efficient DINOv3-ViT-L-Sat-O-Net-BYOL-JSD pipeline reaches 75.33 mDSC / 61.28 mIoU, outperforming standard machine learning baselines fit on the same 24 labelled images with the RGB pixel values used as features. We release CrevasseSeg to support label-efficient segmentation research in remote sensing.

Comments13 pages, 5 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑