arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29081cs.CV

AdapToPASS:面向全景语义分割的歧义感知自适应球形Transformer

AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation

Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

AdapToPASS是一种生物启发的球形Transformer,通过自适应建模歧义提升全景语义分割性能,在未见过的球形变换下较次优方法相对mIoU提升最高18.77%,轻量变体参数不足2M且性能优于紧凑基线。

中文摘要 AI 辅助

球形Transformer通过直接在球形几何上操作并缓解投影引起的失真,已成为全景语义分割(PASS)领域颇具前景的框架。然而,现有架构通常假设球形结构为标准形式且视点稳定,但在实际图像中,由于相机运动不受约束,这些假设常被违背,从而引入上下文与几何歧义。因此,现有架构缺乏处理此类歧义的自适应机制,限制了其对未见过的球形变换的鲁棒性。相比之下,生物感知天生具有歧义感知能力,能适应几何与上下文变化导致的线索可靠性波动,在复杂变换下维持稳定的感知结果。受此启发,我们首先系统分析了现有PASS架构在各种未见过的球形变换下的表现,随后提出AdapToPASS,一种新颖的生物启发式球形Transformer,可自适应建模上下文与几何歧义以实现鲁棒的PASS。其核心是自适应球形注意力(AdaSpA)模块,能根据局部上下文歧义动态调整注意力,模拟生物感知中上下文驱动的自适应机制;为解决几何歧义,AdapToPASS采用双焦点球形表示以平衡视野与空间分辨率,同时引入受生物视觉边界敏感特性启发的边界监督。在室内与室外语义分割任务中,AdapToPASS始终优于现有最先进方法:在未见过的球形变换下,其在Stanford2D3D数据集上的相对mIoU较次优方法提升13.38%,在WildPASS数据集上提升18.77%。我们还推出了参数规模小于2M的轻量变体AdapToPASS-Swift,该变体在保持对球形变换鲁棒性的同时,性能优于紧凑基线方法。

英文摘要

Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which are frequently violated in real-world imagery due to unconstrained camera motion, introducing contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to handle such ambiguity, limiting robustness to unseen spherical transformations. In contrast, biological perception is inherently ambiguity-aware, adapting to fluctuations in cue reliability caused by geometric and contextual variations to maintain stable interpretation under complex transformations. Motivated by this, we first systematically analyze existing PASS architectures under various unseen spherical transformations. We then introduce AdapToPASS, a novel bio-inspired Spherical Transformer that adaptively models contextual and geometric ambiguities for robust PASS. At its core, Adaptive Spherical Attention (AdaSpA) blocks dynamically modulate attention according to local contextual ambiguity, mimicking adaptive, context-driven biological perception. To address geometric ambiguity, AdapToPASS employs Bifocal Spherical Representation to balance field of view and spatial resolution, together with boundary supervision inspired by the boundary-sensitive nature of biological vision. Across indoor and outdoor semantic segmentation, AdapToPASS consistently outperforms prior state-of-the-art methods. Under unseen spherical transformations, it surpasses the next-best method by +13.38% relative mIoU on Stanford2D3D and +18.77% on WildPASS. We further introduce AdapToPASS-Swift, a lightweight variant with fewer than 2M parameters, which surpasses compact baselines while retaining robustness to spherical transformations.

发表机构

  • NTU Singapore(新加坡南洋理工大学)
  • HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑