发表机构
Centre for Data Science and Artificial Intelligence Victoria University of Wellington; Department of Computer Science and Software Engineering University of Canterbury(惠灵顿维多利亚大学数据科学与人工智能中心; 坎特伯雷大学计算机科学与软件工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过固定其他因素仅更换编码器,发现卷积和混合骨干优于Transformer,参数少者表现更好,分割与深度排名一致,部分编码器崩溃为全树分割。
AI 中文摘要
机器人修剪树木需要每个像素的两个事实:它是否属于一棵树,以及它的距离。这两者通常通过附加在视觉骨干网络上的任务头获得,而骨干网络的选择往往基于声誉而非测量。在保持数据集、解码器、损失函数、调度和评估固定的情况下,我们探究:编码器选择对薄植被的联合语义分割和立体深度有多大影响?我们构建了一个硬参数共享网络,一个编码器馈送两个分支,仅交换编码器而不进行下游重新调整。我们在约25M参数预算下,从零训练,评估了[N]个编码器,涵盖[M]个架构家族(CNN、Transformer、混合模型、MLP混合器、状态空间模型)。深度仅在树木像素上评估;分割使用边界F1和背景IoU以防止“将所有内容标记为树”的捷径。三个发现脱颖而出。首先,最强的编码器是卷积和混合模型,而非Transformer:[BestEncoder]以[MIoU]分割mIoU和[Delta]深度δ1领先,而[Y]个普通视觉Transformer中的[X]个在从零训练时崩溃。其次,参数数量不能预测质量——仅[P]M参数的[SmallEncoder]排名超过大两个数量级的模型。第三,分割和深度排名高度一致(Spearman ρ = [RhoValue]),表明没有任务冲突。最后,[N]个编码器中的[K]个崩溃为退化的全树分割——这被边界F1暴露,但被区域IoU隐藏。
英文摘要
A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a vision backbone chosen by reputation rather than measurement. Holding dataset, decoders, losses, schedule, and evaluation fixed, we ask: how much does the encoder choice change joint semantic segmentation and stereo depth on thin vegetation? We build a hard parameter-sharing network with one encoder feeding both branches, swapping only the encoder without downstream retuning. We evaluate [N] encoders across [M] architecture families (CNNs, transformers, hybrids, MLP-mixers, state-space models) near a ~25M budget, trained from scratch. Depth is evaluated on tree pixels only; segmentation uses boundary F1 and background IoU to prevent "label-everything-tree" shortcuts. Three findings stand out. First, the strongest encoders are convolutional and hybrid, not transformers: [BestEncoder] leads with [MIoU] segmentation mIoU and [Delta] depth $δ_1$, while [X] of [Y] plain vision transformers collapse when trained from scratch. Second, parameter count does not predict quality --- [SmallEncoder] at only [P]M parameters outranks models two orders of magnitude larger. Third, segmentation and depth rankings agree strongly (Spearman $ρ$ = [RhoValue]), showing no task conflict. Finally, [K] of [N] encoders collapse to degenerate all-tree segmentation --- exposed by boundary F1 but hidden by region IoU.