发表机构
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院电气工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人感知系统多任务学习中现有解码策略瓶颈,提出DPNeXt框架,用双深度可分离倒置瓶颈及MTBG策略,在多任务密集预测上表现出色,相比DPT大幅减少参数并提升推理速度。
AI 中文摘要
机器人感知系统中的多任务学习(MTL)通过整合语义分割和深度估计来支持对3D空间场景的全面理解。虽然视觉基础模型(VFM)越来越多地被用作强大的特征编码器,但现有的解码策略是一个关键瓶颈。为解决此问题,我们提出了DPNeXt,这是一种简化的多尺度特征融合解码器,是标准密集预测Transformer(DPT)的有效替代方案。DPNeXt使用双深度可分离倒置瓶颈,通过以融合为中心的解码和独立任务模块化来提高冻结VFM的利用率。为进一步减轻任务之间的负归纳转移,我们引入了多任务边界引导(MTBG)策略。与添加融合模块或门控的先前边界感知方法不同,MTBG应用对称的以边界为重点的监督来鼓励几何一致性,而无需额外的注释或推理成本。在Cityscapes上的实验表明,DPNeXt-S优于先前的最新(SOTA)MTL模型,而DPNeXt-B进一步提高了整体性能,并在比较方法中取得了最佳结果。在NYUv2上,DPNeXt-B在比较方法中也取得了最佳的语义分割和深度估计结果,同时所需的可训练参数比先前的大规模MTL模型少得多。与标准DPT相比,DPNeXt-S减少了78.6%的可训练参数,并在资源受限的笔记本电脑硬件上的比较模型中实现了最快的推理速度。源代码、模型检查点和演示视频将在这个https URL上提供。
英文摘要
Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the standard Dense Prediction Transformer (DPT). DPNeXt uses dual depthwise separable inverted bottlenecks to improve frozen VFM utilization through fusion-centric decoding and independent task modularization. To further mitigate negative inductive transfer between tasks, we introduce the Multi-Task Boundary Guidance (MTBG) strategy. Unlike prior boundary-aware methods that add fusion modules or gating, MTBG applies symmetric boundary-focused supervision to encourage geometric consistency without extra annotation or inference cost. Experiments on Cityscapes show that DPNeXt-S outperforms prior state-of-the-art (SOTA) MTL models, while DPNeXt-B further improves the overall performance and achieves the best results among the compared methods. On NYUv2, DPNeXt-B also achieves the best semantic segmentation and depth estimation results among the compared methods while requiring substantially fewer trainable parameters than prior large-scale MTL models. Compared with the standard DPT, DPNeXt-S reduces trainable parameters by 78.6% and achieves the fastest inference speed among the compared models on resource-constrained laptop hardware. The source code, model checkpoints, and a demo video will be made available at https://github.com/kangjehun/DPNeXt.
Comments8 pages, 5 figures. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)