发表机构
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对城市环境语义映射中尺度与观测距离不匹配及异构映射策略难题,提出类别感知校准策略、互信息正则化和类别级价值估计器,并引入大规模数据集,通过帕累托优化实现多尺度可靠映射。
AI 中文摘要
语义映射是具身导航的基础,然而现有方法主要针对室内环境开发,其中物体的尺度变化相对有限,且观测视角范围受限。城市环境带来了更大的挑战:智能体必须在具有高度多样化观看距离的大型空间中导航,同时映射从行人到建筑物等不同尺度的物体。这些条件引入了现有数据集和方法未能覆盖的两个关键难点。首先,物体尺度与观测距离可能严重不匹配。例如,小物体可能被远距离观看,而大物体可能在极近距离被观测,导致观测似然不可靠。其次,尺寸和几何形状差异显著的物体需要不同的映射行为,这难以用单一共享的价值估计器捕获。为研究这些挑战,我们引入了一个大规模城市语义映射数据集,具有逼真的城市布局、高保真渲染以及覆盖多种物体尺度的实例级标注。然后,我们提出一种类别感知的似然校准策略,根据物体类别和观看距离识别并缓解不可靠的观测。由于校准策略和运动策略针对同一映射目标进行优化,它们可能学习冗余的捷径并变得过度耦合。因此,我们引入一个互信息(MI)正则化器,惩罚它们估计的表示依赖性,并鼓励互补行为。为了更好地建模跨物体尺度的异构映射策略,我们进一步采用类别级价值估计器。我们将它们的联合优化表述为帕累托优化问题,以缓解跨类别的梯度冲突。
英文摘要
Semantic mapping is fundamental to embodied navigation, yet existing methods are developed for indoor environments, where objects exhibit relatively limited scale variation and are observed from a restricted range of viewpoints. Urban environments pose substantially greater challenges: agents must map objects ranging from pedestrians to buildings while navigating large spaces with highly diverse viewing distances. These conditions introduce two key difficulties that existing datasets and methods fail to cover. First, object scale and observation distance can be severely mismatched. For example, small objects may be viewed from far away, whereas large objects may be observed at extremely close range, resulting in unreliable observation likelihoods. Second, objects with substantially different sizes and geometries require distinct mapping behaviors, which are difficult to capture with a single shared value estimator. To investigate these challenges, we introduce a large-scale urban semantic mapping dataset featuring realistic city layouts, high-fidelity rendering, and instance-level annotations spanning multiple object scales. We then propose a category-aware likelihood calibration policy that identifies and alleviates unreliable observations according to object category and viewing distance. Because the calibration and motion policies are optimized toward the same mapping objective, they may learn redundant shortcuts and become excessively coupled. We therefore introduce a mutual-information (MI) regularizer that penalizes their estimated representation dependence and encourages complementary behaviors. To better model heterogeneous mapping strategies across object scales, we further employ category-wise value estimators. We formulate their joint optimization as a Pareto optimization problem to mitigate conflicting gradients across categories.