ST-LoRA:用于不确定性感知农业分割的单轨迹LoRA集成
ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation
浏览论文内容
中文总结 AI 辅助
ST-LoRA结合LoRA与快照集成,从单轨迹构建高效集成,在农业分割任务中,其性能优于多种基线方法,能高效实现不确定性感知。
中文摘要 AI 辅助
数字农业中可靠的决策支持需要准确的预测和校准良好的不确定性估计,尤其是对于语义分割这类密集预测任务。集成方法能提供强大的不确定性量化能力,但其计算和内存需求限制了实际应用,而单模型近似方法往往会在效率上牺牲不确定性质量。我们提出ST-LoRA,这是一种参数高效的集成框架,通过将低秩适应(LoRA)与快照集成(snapshot ensembling)相结合,从单一训练轨迹构建多样化的集成成员。每个成员共享一个冻结的预训练主干网络,仅在轻量级低秩适配器上存在差异,可将可训练参数减少至完整模型的10%以下,同时保持集成的多样性。我们在两个农业数据集上进行评估:GrowliFlower-L(露天种植的花椰菜)和BUP20(温室种植的甜椒),使用SegFormer和Mask2Former两种模型,覆盖分布内性能、分布偏移下的校准以及分布外(OoD)检测。消融实验表明,对于密集预测,前馈层而非注意力层是LoRA的关键目标,这与语言模型中仅针对注意力层的惯例相反。ST-LoRA在两个数据集和两种模型上的分割准确率和校准性能与全秩集成相当或更优,同时大幅减少了训练时间、推理延迟、内存占用和存储需求。与高效基线方法——快照集成、蒙特卡洛(MC)Dropout以及深度确定性不确定性(Deep Deterministic Uncertainty)相比,ST-LoRA在图像/像素级OoD检测、偏移下的校准稳定性以及跨种子方差方面始终与这些方法相当或更优,且参数更少、计算量更低。这些结果表明,LoRA高效集成适应是构建不确定性感知农业视觉系统的一种高度有效且实用的方法。
英文摘要
Reliable decision support in digital agriculture requires not only accurate predictions but also well-calibrated uncertainty estimates, particularly for dense prediction tasks such as semantic segmentation. Ensembles provide strong uncertainty quantification but are computationally and memory demanding, while single-model approximations often sacrifice uncertainty quality. We propose ST-LoRA, a parameter-efficient ensemble that builds diverse members from a single training trajectory by combining Low-Rank Adaptation (LoRA) with snapshot ensembling. All members share a frozen pretrained backbone and differ only in lightweight low-rank adapters, which sharply reduces trainable parameters, checkpoint storage, and I/O overhead. We evaluate SegFormer, Mask2Former, and EoMT on GrowliFlower-L (cauliflower, open field) and BUP20 (sweet pepper, glasshouse), covering in-distribution performance, calibration under covariate shift, and near- and far-out-of-distribution (OoD) detection, with BUTom21 (tomato) as near-OoD data. Extensive ablations show that feed-forward layers, not attention projections, are the critical LoRA target for dense prediction, and that the scaling ratio $α/r$ governs an accuracy--calibration trade-off. Against full-rank snapshot ensembles, ST-LoRA is competitive in segmentation quality, with architecture-dependent training time and energy savings. Against MC Dropout, DDU, and six post-hoc calibrators, it achieves the strongest far-OoD image-level detection and near-OoD pixel-level localization with low cross-seed variance, although full-rank ensembles remain better calibrated. These results show that LoRA-based ensembling offers a compelling efficiency--performance trade-off for agricultural vision systems.
发表机构
- University of Bonn(波恩大学)
- Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn(波恩大学拉马尔机器学习与人工智能研究所)
- Commonwealth Scientific and Industrial Research Organisation (CSIRO)(英联邦科学与工业研究组织(CSIRO))
机构由 AI 辅助整理,请以论文原文为准。