发表机构
Heinrich-Heine-University Düsseldorf; Hamad Bin Khalifa University (HBKU); Linköping University(海因里希·海涅大学杜塞尔多夫; 哈马德·本·哈利法大学; 林雪平大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于Wasserstein目标的证据深度学习框架,用于语义分割中的分布外检测,在多个基准上验证了最优阶数依赖架构,并仅需单次前向传播即可提升性能。
AI 中文摘要
语义分割网络在固定的类别集合上运行,因此在部署期间出现分布外(OOD)对象时会失效,这是自动驾驶等安全关键应用中的一个关键限制。可靠地识别OOD对象需要校准良好的认知不确定性,然而常见的基于softmax的置信度分数仍然过于自信,而贝叶斯替代方法如蒙特卡洛dropout或深度集成则需要昂贵的重复前向传播。证据深度学习(EDL)通过将类别概率建模为从单个确定性前向传播中学习的Dirichlet分布,提供了一种高效的替代方案。现有的EDL公式依赖于欧几里得目标,将预测推向单纯形的顶点,鼓励过度自信而非保留对不熟悉输入的不确定性。我们转而采用基于Wasserstein的目标,这些目标尊重概率单纯形的几何结构,并在统一的证据框架内研究Wasserstein阶数对分割准确性和OOD检测的影响。我们在SegmentMeIfYouCan基准上,对卷积(DeepLabV3+)和基于Transformer(SegFormer)的架构评估了该框架,包括LostAndFound、RoadObstacle21、RoadAnomaly21和Fishyscapes。我们的结果表明,最优的Wasserstein阶数依赖于架构:二阶目标在卷积骨干上占优,三阶目标在Transformer骨干上占优,并且我们的框架在大多数指标上超越了可比基线,且仅需单个确定性前向传播。
英文摘要
Semantic segmentation networks operate on a fixed set of classes and therefore fail when out-of-distribution (OOD) objects appear during deployment, a critical limitation for safety-critical applications such as autonomous driving. Reliably identifying OOD objects requires well-calibrated epistemic uncertainty, yet common softmax-based confidence scores remain overconfident, while Bayesian alternatives such as Monte Carlo dropout or deep ensembles require costly repeated forward passes. Evidential Deep Learning (EDL) offers an efficient alternative by modeling class probabilities as a Dirichlet distribution learned from a single deterministic forward pass. Existing EDL formulations rely on Euclidean objectives that push predictions towards the simplex vertices, encouraging overconfidence rather than preserving uncertainty for unfamiliar inputs. We instead employ Wasserstein-based objectives, which respect the geometry of the probability simplex, and study the influence of the Wasserstein order on segmentation accuracy and OOD detection within a unified evidential framework. We evaluate this framework on a convolutional (DeepLabV3+) and a transformer-based (SegFormer) architecture on the SegmentMeIfYouCan benchmark, including LostAndFound, RoadObstacle21, RoadAnomaly21, and Fishyscapes. Our results show the optimal Wasserstein order is architecture-dependent: second-order objectives dominate on the convolutional backbone, third-order objectives on the transformer backbone, and our framework surpasses comparable baselines on most metrics, with a single deterministic forward pass.
Comments20 pages, 6 images, 3 figures