发表机构
ETH Zürich; INSAIT, Sofia University; University of Bonn; Microsoft(苏黎世联邦理工学院; 保加利亚索菲亚大学INSAIT研究所; 波恩大学; 微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DVPSFormer是一种统一在线架构,通过ESD和OMV机制实现高效4D场景理解,在两个DVPS基准上达SOTA,为自动驾驶提供在线机器人感知的精简方案。
AI 中文摘要
安全的自主导航需要对动态环境进行整体理解,需同时估计度量深度、语义分割和实例轨迹。深度感知视频全景分割(DVPS)统一了这些任务,但现有方法常依赖计算成本高昂的多阶段流水线或离线跟踪,不适合实时决策。为解决该问题,我们提出DVPSFormer,一种用于高效4D场景理解的统一在线架构。我们方法的核心是显式场景离散化(ESD),这是一种利用分割查询表示前景和背景区域的新机制,使离散到连续(D2C)深度头能单次解码度量深度,该机制紧密耦合语义与几何学习,同时显著降低延迟。此外,我们提出在线多数投票(OMV)机制,利用时间一致性在实例跟踪期间优化分类。DVPSFormer在Cityscapes-DVPS和SemKITTI-DVPS基准上取得新的 state-of-the-art,为在线机器人感知提供了精简解决方案。代码和模型可在this https URL获取。
英文摘要
Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semantic segmentation, and instance trajectories. While depth-aware video panoptic segmentation (DVPS) unifies these tasks, existing approaches often rely on computationally expensive, multi-stage pipelines or offline tracking, rendering them unsuitable for real-time decision-making. To address this, we propose DVPSFormer, a unified online architecture designed for efficient 4D scene understanding. Central to our approach is explicit scene discretization (ESD), a novel mechanism that leverages segmentation queries to represent foreground and background regions, enabling a discrete-to-continuous (D2C) depth head to decode metric depth in a single pass. This tightly couples semantic and geometric learning while significantly reducing latency. Furthermore, we propose an online majority voting (OMV) mechanism that exploits temporal consistency to refine classification during instance tracking. DVPSFormer establishes a new state-of-the-art on the Cityscapes-DVPS and SemKITTI-DVPS benchmarks, offering a streamlined solution for online robotic perception. Code and models are available at https://royyang0714.github.io/DVPSFormer.