发表机构
Technical University of Munich; Huawei Hilbert Research Center (Dresden)(慕尼黑工业大学; 华为希尔伯特研究中心(德累斯顿))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VideoReloc利用自适应视频片段和语义场景图,通过假设优先注册与方向细化,实现光照和家具变化下的长时室内重定位,以100 kB地图达到高成功率。
AI 中文摘要
给定一个紧凑的语义场景图,长时室内视频重定位在光照和家具变化后估计地图坐标系下的轨迹。视觉方法依赖外观,在这些变化下变得不可靠;而逐帧仅利用物体类别和几何进行定位,则会留下稀疏且模糊的证据。我们提出了VideoReloc,其自适应片段利用里程计收集空间证据,直到满足物体和运动标准,从而根据观察到的场景调整查询长度。其运行级决策利用跨连接片段累积的证据重新检查冲突的放置,稳定了超越相邻片段跟踪的轨迹。假设优先的注册方法从物体三元组提出位姿,并使用片段范围的物体中心和盒子表面验证每个位姿。面向方向的细化利用盒子面、重力和墙壁方向来解决相机方向的模糊性并细化完整位姿。这将稀疏地图重定位重新定义为对空间扩展视频查询的验证,将判别性支持从存储的外观转移到时间上下文,并允许使用100 kB的类别标记盒子地图。在RIO10和ReplicaCAD上,因果评估下1米/10度误差内的全帧定位成功率为73.5%和61.1%,在片段闭合后提升至90.6%和74.8%。评估的逐帧场景坐标回归器分别达到47.6%和49.8%,地图大小为12.6-42 MB。项目页面:此https URL
英文摘要
Given a compact semantic scene graph, long-term indoor video relocalization estimates a map-frame trajectory after lighting and furniture changes. Visual methods rely on appearance and become unreliable under these changes; localizing one frame at a time from object classes and geometry instead leaves sparse, ambiguous evidence. We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene. Its run-level decision rechecks conflicting placements using evidence accumulated across connected clips, stabilizing the trajectory beyond adjacent-clip tracking. Hypothesis-first registration proposes poses from object triplets and verifies each using clip-wide object centers and box surfaces. Orientation-aware refinement uses box faces, gravity and wall directions to resolve ambiguity in camera orientation and refine the full pose. This reframes sparse-map relocalization as verification of spatially extended video queries, moving discriminative support from stored appearance to temporal context and permitting a 100 kB map of class-labelled boxes. On RIO10 and ReplicaCAD, the all-frame localization success rate at 1 m/10$^\circ$ is 73.5% and 61.1% under causal evaluation, rising to 90.6% and 74.8% with clip closure. The evaluated per-frame scene coordinate regressors reach up to 47.6% and 49.8%, respectively, with maps of 12.6-42 MB. Project page: https://videoreloc.github.io
Comments8 pages, 3 figures, 4 tables. Project page: https://videoreloc.github.io