arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在共性期望最大化用于框监督的树冠实例分割

Latent Commonality Expectation-Maximisation for Box-supervised Tree Crown Instance Segmentation

Thomas Pitts, Kunqi Li, Bin Liang

arXiv 2609.26549首次发表:更新:

发表机构

University of Technology Sydney(悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LACE框监督分割模型,利用冻结自监督特征和期望最大化分离树冠共性,仅用框标注即超越掩膜监督基线,适用于稀疏冠层环境。

AI 中文摘要

从航空影像中进行单木树冠分割是景观尺度上树木级碳核算、生物多样性和恢复监测的基础。然而,现有模型主要针对密集冠层森林影像训练,在稀树草原和干旱地区表现退化,这些地区的树冠稀疏、外观多变,且在标注基准中代表性不足。这些模型通常还依赖于昂贵的多边形标注。我们引入了LACE(潜在共性期望最大化),一种框监督的实例分割模型,在0.1米/像素的航空RGB树冠影像上进行了评估。LACE使用冻结的DINOv3-web ViT-L/16编码器,在四个空间偏移处应用并交错成更密集的特征网格,配以轻量级CenterNet风格的检测头,仅使用边界框进行训练。我们利用期望最大化将边界框内反复出现的外观(即“树性”)与周围环境分离。在OAM-TCD基准测试集上,LACE在900张框标注图像上训练且无掩膜标注的情况下,达到了掩膜AP$_{50}$为$0.663 \pm 0.001$(3个随机种子),高于Restor发布的掩膜监督Mask R-CNN所取得的0.626,后者在完整的约4.2k图像集上训练。在稀疏冠层保留集上,掩膜AP$_{50}$升至0.691,而掩膜监督基线Detectree2为0.612。在NeonTreeEvaluation上,使用官方评估代码,LACE仅从23,424个手工标注的RGB框达到$0.728 \pm 0.003$的F1@0.4(5个随机种子),与作者DeepForest模型公布的0.719相当,而使用的训练标注不足其0.1%,且未使用其LiDAR衍生的3000万树冠预训练集。通过利用冻结的自监督特征,LACE仅从框标注即可匹配或超越全监督专家基线,消除了在标注数据稀缺的稀疏冠层环境中进行树冠实例分割时对多边形标注的需求。

英文摘要

Individual tree crown segmentation from aerial imagery underpins tree-level carbon accounting, biodiversity, and restoration monitoring at landscape scale. However, existing models are predominantly trained on dense canopy forest imagery and degrade in savannah and drylands, where tree crowns are sparse, of variable appearance, and underrepresented in annotated benchmarks. These models also typically depend on costly polygon annotations. We introduce LACE (LAtent Commonality Expectation-maximisation), a box-supervised instance segmentation model, evaluated on 0.1 m/px aerial RGB tree crown imagery. LACE uses a frozen DINOv3-web ViT-L/16 encoder, applied at four spatial offsets and interlaced into a denser feature grid, with a lightweight CenterNet-style detection head trained solely on bounding boxes. We use expectation-maximisation to separate recurring appearance, the "treeness", within bounding boxes from surroundings. On the OAM-TCD benchmark test set, LACE reaches a mask AP$_{50}$ of $0.663 \pm 0.001$ (3 seeds) trained on 900 box-annotated images and without mask annotations, above the 0.626 scored by Restor's released mask-supervised Mask R-CNN, which was trained on the full ~4.2k image set. On a sparse-canopy holdout set, mask AP$_{50}$ rises to $0.691$ versus $0.612$ for Detectree2, a mask-supervised baseline. On NeonTreeEvaluation, using the official evaluation code, LACE reaches $0.728 \pm 0.003$ F1@0.4 (5 seeds) from 23,424 hand-annotated RGB boxes alone, matching the authors' DeepForest model's published 0.719, using under 0.1% of its training annotations and none of its LiDAR-derived 30M-crown pretraining set. By leveraging frozen self-supervised features, LACE matches or surpasses fully-supervised specialist baselines from boxes alone, removing the need for polygon annotation in tree crown instance segmentation for sparse-canopy environments where labelled data is scarce.

Comments37 pages, 18 tables. Code and model checkpoints to be released upon publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑