TDFNet:用于全景显著目标检测的三投影可变形融合网络
TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection
浏览论文内容
中文总结 AI 辅助
针对全景显著目标检测中投影带来的几何畸变问题,提出首个三投影可变形融合网络TDFNet,通过跨投影可变形注意力模块与纬度引导融合模块,实现特征优化,提升检测性能。
中文摘要 AI 辅助
近年来,全景显著目标检测在机器人视觉、虚拟现实及相关应用中展现出日益增长的潜力。然而,将球面场景投影到二维平面不可避免地会引入几何畸变,从根本上限制了现有基于投影的方法的有效性。具体而言,等距圆柱投影(ERP)会遭受严重的极地拉伸畸变,而立方体贴图投影则会在立方体面边界处引入不连续性,导致特征判别能力下降和几何一致性受损。为解决这些局限,我们提出TDFNet,这是首个用于全景显著目标检测的三投影可变形融合网络,利用互补的投影表示来缓解几何畸变并提升检测性能。首先,我们设计了跨投影可变形注意力(CDA)模块,该模块利用不同投影之间的空间对应关系来构建几何感知的采样位置,指导可变形注意力进行跨投影上下文聚合,增强对投影诱导畸变的鲁棒性。此外,我们引入了纬度引导融合(LGF)模块,该模块利用球面纬度先验构建几何置信权重,以自适应平衡ERP和立方体贴图投影(CMP)的特征;同时,LGF融入了来自正形投影(Tangent Projection)的低畸变语义参考,实现跨投影特征细化和空间一致性增强。通过构建基于ERP、CMP和正形投影的三分支编码架构,TDFNet同时保留了全局空间连续性、局部几何细节和细粒度边界信息。
英文摘要
Recent years have witnessed the growing potential of panoramic salient object detection in robotic vision, virtual reality, and related applications. However, projecting spherical scenes onto 2D planes inevitably introduces geometric distortions, which fundamentally limit the effectiveness of existing projection-based methods. Specifically, Equirectangular Projection (ERP) suffers from severe polar stretching distortions, while cube map projection introduces discontinuities across cube-face boundaries, resulting in degraded feature discriminability and compromised geometric consistency. To address these limitations, we propose TDFNet, the first Tri-projection Deformable Fusion Network for panoramic salient object detection, exploiting complementary projection representations to alleviate geometric distortions and improve detection performance.Specifically, we design a cross-projection deformable attention (CDA) module that leverages spatial correspondences between different projections to construct geometry-aware sampling locations, guiding deformable attention for cross-projection contextual aggregation and enhancing robustness against projection-induced deformations. Furthermore, we introduce a latitude-guided fusion module, which utilizes spherical latitude priors to construct geometric confidence weights for adaptively balancing ERP and CMP features. Meanwhile, LGF incorporates distortion-reduced semantic references from Tangent Projection to achieve cross-projection feature refinement and spatial alignment.By constructing a three-branch encoding architecture based on ERP, CMP, and Tangent Projection, TDFNet simultaneously preserves global spatial continuity, local geometric details, and fine-grained boundary information.
发表机构
- School of Artificial Intelligence, Jiangxi Normal University(江西师范大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。