发表机构
College of Electronics and Information Engineering, Shenzhen University; Guangdong Key Laboratory of Intelligent Information Processing(深圳大学电子与信息工程学院; 广东省智能信息处理重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对室内语义占用预测的长尾类别瓶颈,提出Group-UFD Occ方法,通过分组策略与UFD损失优化,在EmbodiedScan数据集上较基线实现11.38%相对提升,长尾类别精度显著提高。
AI 中文摘要
近年来,3D语义占用预测因用于理解室内场景而受到越来越多关注。但与结构化的室外环境不同,室内场景具有类别多样性高且呈现严重长尾分布的特点,这已成为限制现有模型性能的核心瓶颈。为解决该挑战,我们提出一种基于分层语义监督与协同损失优化的新方法Group-UFD Occ。在架构层面,我们引入细粒度语义分组策略,并设计多尺度并行的“主专家”预测头,通过深度正则化引导模型高效学习尾类特征;在优化层面,我们引入统一焦点-骰子(Unified Focal-Dice, UFD)损失,该协同损失函数会在每个体素层面动态聚焦难样本,同时从基于区域的视角同步优化预测物体的几何完整性。我们在大规模EmbodiedScan数据集上开展实验,结果表明,与基线相比,该方法实现了11.38%的相对提升,且在多个关键长尾类别上取得了显著的精度增益。
英文摘要
Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor environments, indoor scenes feature a high diversity of object categories that exhibit a severe long-tailed distribution, which has become a core bottleneck limiting the performance of existing models. To tackle this challenge, we propose a novel method, Group-UFD Occ, based on hierarchical semantic supervision and synergistic loss optimization. At the architectural level, we introduce a fine-grained semantic grouping strategy and design multi-scale, parallel ``main-expert'' prediction heads to guide the model in efficiently learning tail-class features through deep regularization. At the optimization level, we introduce the Unified Focal-Dice (UFD) loss. This synergistic loss function dynamically focuses on hard samples at the per-voxel level. Meanwhile, it simultaneously optimizes the geometric integrity of predicted objects from a region-based perspective. We conducted experiments on the large-scale EmbodiedScan dataset. The results demonstrate that our method yields a relative improvement of 11.38\% over the baseline, with substantial accuracy gains in several critical long-tailed categories.
Comments8 pages, 2 figures