发表机构
Tongji University; Loughborough University(同济大学; 拉夫堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对拥挤室内环境中四足机器人单目语义场景补全受遮挡影响的问题,提出CrowdOcc数据集与框架,通过法线引导融合和以人为中心的稀疏交互,在测试集上达到15.80 IoU等最先进性能。
AI 中文摘要
在真实拥挤的室内环境中,四足机器人的单目语义场景补全(SSC)仍未得到充分探索,其中人-场景遮挡会破坏静态几何结构,且人体占用预测往往不完整或空间错位。我们提出了CrowdOcc,一个针对该场景的RGB-D数据集和单目SSC框架。CrowdOcc包含来自11个室内场景的25.1K帧,通过静态-动态解耦构建语义占用标注。我们的框架结合了:(i)法线引导的场景几何融合(NGSGF),利用表面法线线索补充深度感知的提升,以实现对遮挡鲁棒的几何重建;以及(ii)以人为中心的稀疏交互(HCSI),在3D空间中选择性地建模人与人以及局部人-场景关系。我们的方法在CrowdOcc的场景分离测试集上达到了最先进的SSC性能,分别取得15.80 IoU、11.40 mIoU和46.23 Human IoU,展示了其对未见室内场景的泛化能力。
英文摘要
Monocular semantic scene completion (SSC) for quadruped robots remains underexplored in real crowded indoor environments, where human-scene occlusion disrupts static geometry and human occupancy predictions are often incomplete or spatially misplaced. We present CrowdOcc, an RGB-D dataset and monocular SSC framework for this setting. CrowdOcc contains 25.1K frames from 11 indoor scenes, with semantic occupancy annotations constructed through static dynamic decoupling. Our framework combines: (i) Normal Guided Scene Geometry Fusion (NGSGF) to complement depth-aware lifting with surface-normal cues for occlusion robust geometry; and (ii) Human-Centric Sparse Interaction (HCSI) to selectively model human-human and local human scene relations in 3D. Our method achieves state-of-the-art SSC performance on CrowdOcc's scene-disjoint test set, reaching 15.80 IoU, 11.40 mIoU, and 46.23 Human IoU, demonstrating generalization to unseen indoor scenes.
Comments8 pages, 4 figures. Submitted to IEEE International Conference on Robotics and Automation (ICRA) 2027