探索二维骨干网络对室内语义占用预测的影响
Exploring 2D backbone effects for indoor semantic occupancy prediction
- College of Electronics and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过替换图像骨干网络,发现DINOv2等更强编码器显著提升室内语义占用预测的mIoU,表明骨干网络选择是影响三维预测的关键因素。
AI中文摘要:
语义占用预测为具身智能体提供了体素级别的空间描述,表明空间哪些区域是空闲的、被占用的以及具有语义意义。在诸如EmbodiedScan等RGB-D处理流程中,图像编码器通常作为默认模块保留,尽管其提取的特征是后续采样到三维网格中的视觉证据。我们直接研究了这一设计选择。一个核心发现是,更换二维骨干网络比精心设计的几种占用预测架构或模块更能提升占用预测的准确性。我们保持主要的RGB-D投影、深度分支和占用预测头不变,仅替换图像骨干网络。比较的编码器包括CLIP-ResNet、CLIP-ViT、BLIP2和DINOv2。在受控设置下,测得的mIoU变化显著:DINOv2达到30.55%,BLIP2达到29.49%,CLIP-ViT达到24.33%,CLIP-ResNet达到17.41%。更强的编码器在不修改下游三维融合流程的情况下,也超过了原始的EmbodiedScan ResNet-50基线。类别级结果提供了更详细的图景:DINOv2在许多布局和结构类别上表现更强,而BLIP2在若干以对象为中心的类别上保持接近。CLIP-ViT明显优于CLIP-ResNet,表明CLIP特征作为密集令牌的暴露方式对体素提升至关重要。这些结果表明,在具身语义占用预测中,图像骨干网络并非次要的工程细节,而是最终三维预测变异的主要来源。
英文摘要:
Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a default module, even though its features are the visual evidence later sampled into the 3D grid. We study this design choice directly. A central finding is that changing the 2D backbone improves occupancy accuracy more than several carefully designed occupancy architectures or modules. We keep the main RGB-D projection, depth branch, and occupancy head fixed, and replace only the image backbone. The compared encoders are CLIP-ResNet, CLIP-ViT, BLIP2, and DINOv2. Under the controlled setting, the measured mIoU changes substantially: DINOv2 obtains 30.55\%, BLIP2 obtains 29.49\%, CLIP-ViT obtains 24.33\%, and CLIP-ResNet obtains 17.41\%. The stronger encoders also exceed the original EmbodiedScan ResNet-50 baseline without modifying the downstream 3D fusion pipeline. Class-level results give a more detailed picture: DINOv2 is stronger on many layout and structural categories, whereas BLIP2 remains close on several object-centered classes. CLIP-ViT improves clearly over CLIP-ResNet, showing that the way CLIP features are exposed as dense tokens matters for voxel lifting. These results indicate that the image backbone is not a secondary engineering detail in embodied semantic occupancy, but a major source of variation in the final 3D prediction.