ProClosure:基于渐进边界闭合的单目视频层次化房间-物体分配
ProClosure: Hierarchical Room-Object Assignment using Progressive Boundary Closure from Monocular Video
浏览论文内容
中文总结 AI 辅助
提出渐进边界闭合方法,从单目视频恢复房间层级并分配物体,通过渐进加厚边界闭合开口,利用相机位姿种子实现确定性分割,显著提升房间和物体分配精度。
中文摘要 AI 辅助
3D场景图将物体分组到房间中。当机器人被要求从厨房取一个物体时,正是这种分组告诉它去哪里寻找。记录在错误房间中的物体无法通过查询正确房间名称来检索。我们提出了渐进边界闭合(Progressive Boundary Closure),该方法从单目RGB视频中恢复房间层级。SLAM前端和开放词汇分割器提供结构点云、相机轨迹和物体轨迹。点云被栅格化为俯视地图,从中恢复房间,每个物体归属于占据其大部分范围的房间。难点在于地图本身。墙壁仅在相机观察过的地方被记录,因此边界上的间隙可能是门洞或从未被观察到的墙段;两者无法区分。先前的方法将两者都视为通道,合并了本应保持分离的房间。我们观察到两者需要相同的处理:房间不应跨越其中任何一个,因此两者都被闭合,无需区分。这样的开口在少量边界增长下闭合,且很少有视线穿过它,因此不同房间中的点很少能看到彼此。我们利用前者恢复房间,利用后者将物体分配到房间。通过向内渐进加厚边界,并在每个自由空间区域被包围时将其冻结,从而获得房间,因此每个开口在其自身尺度上密封,而不是在预先固定的半径上密封。相机位姿被用作种子,这消除了采样启发式,并使分割具有确定性。在6个HM3D-Semantics场景的10个楼层中,与HOV-SG在相同的俯视地图上评分,我们为72个标注区域恢复了74个房间(HOV-SG:44个),在IoU 0.25下将房间F1从0.741提高到0.890,但精度有所损失,物体到房间的ARI从0.488提高到0.696(p=0.002,在每个楼层上均领先)。
英文摘要
A 3D scene graph groups objects into rooms. When a robot is asked to fetch an object from the kitchen, that grouping is what tells it where to look. An object recorded in the wrong room is not retrievable by a query naming the correct room. We introduce Progressive Boundary Closure, which recovers room layer from a monocular RGB video. A SLAM front end and an open-vocabulary segmenter supply a structural point cloud, camera trajectory and object tracks. The cloud is rasterised into a top-down map, rooms are recovered from it, and each object takes the room holding most of its extent. The difficulty lies in the map itself. Walls are recorded only where the camera looked, so a gap in the boundary may be a doorway or a stretch of wall that was never observed; nothing distinguishes the two. Prior methods treat both as passages, merging rooms that should remain separate. We observe that both require the same treatment: a room should not extend across either, so both are closed and need not be distinguished. Such an opening closes under a small amount of boundary growth, and few sightlines cross it, so points in different rooms rarely see one another. We use the first to recover rooms and the second to assign objects to them. Rooms are obtained by Progressively thickening the boundary inward and freezing each free-space region once it becomes enclosed, so every opening seals at its own scale rather than at a radius fixed in advance. Camera poses are used as seeds, which removes the sampling heuristic and makes the segmentation deterministic. Over 10 floors of 6 HM3D-Semantics scenes, scored against HOV-SG on identical top-down maps, we recover 74 rooms for 72 annotated regions (HOV-SG: 44), raising room F_1 from 0.741 to 0.890 at IoU 0.25 at some cost in precision, and object-to-room ARI from 0.488 to 0.696 (p=0.002, ahead on every floor).
发表机构
- IIT Jodhpur(印度理工学院焦特布尔分校)
机构由 AI 辅助整理,请以论文原文为准。