arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于联合位置嵌入和遮挡级别注意力的遮挡感知全景分割

Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention

Wenbo Wei, Jun Wang, Shan Raza, Abhir Bhalerao

arXiv 2607.18112首次发表:更新:

AI 中文总结

针对复杂场景全景分割因遮挡面临挑战且现代方法忽视遮挡建模的问题,提出PEMOLA模块,通过联合位置嵌入和遮挡级别注意力改进遮挡建模,实验表明该方法能在低计算开销下提升全景分割质量。

AI 中文摘要

复杂场景中的全景分割因遮挡而具有挑战性,但现代方法往往忽视遮挡建模。本文提出了带有遮挡级别注意力的位置嵌入调制(PEMOLA),这是一种可无缝集成到基于Transformer的全景分割中的新型遮挡感知模块。通过在COCO - OLAC数据集上训练遮挡分类器来获取遮挡级别注意力作为空间指导,并将遮挡标签编码为可学习嵌入以产生通道权重。通过联合调制,PEMOLA将遮挡先验优雅地引入位置嵌入,从而改进遮挡建模。进一步标注了Cityscapes数据集为Cityscapes - OLAC来评估PEMOLA的跨数据集泛化能力。在COCO - OLAC和Cityscapes - OLAC上的大量实验表明,PEMOLA在引入最小计算开销的同时持续提高全景分割质量,突出了遮挡建模的重要性。

英文摘要

Panoptic segmentation in complex scenes remains challenging because of occlusions, yet modern approaches often neglect occlusion modelling. In this paper, we propose Position Embedding Modulation with Occlusion Level Attention (PEMOLA), a novel occlusion-aware module that can be seamlessly integrated into transformer-based panoptic segmentation. To obtain occlusion cues, we train an occlusion classifier on the COCO-OLAC dataset. The classifier derives the occlusion-level attention, which serves as spatial guidance, while the occlusion labels are encoded into a learnable embedding to produce channel-wise weights. Through joint modulation, PEMOLA elegantly introduces the occlusion priors into the position embedding, thereby improving the occlusion modelling. We further annotate the Cityscapes dataset with occlusion levels, termed Cityscapes Occlusion Labels for All Computer Vision Tasks (Cityscapes-OLAC), following the same labelling protocol as COCO-OLAC, to evaluate the cross-dataset generalisation ability of PEMOLA. Extensive experiments on COCO-OLAC and Cityscapes-OLAC demonstrate that PEMOLA consistently improves panoptic segmentation quality while introducing minimal computational overhead. These results highlight the importance of occlusion modelling, where incorporating occlusion-level attention helps deliver robust panoptic segmentation under occlusion. Code and dataset are available at https://github.com/wenbo-wei/PEMOLA.

CommentsAccepted at the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑