Know-Your-Scene (KYS)-SLAM:立体视觉SLAM中用于特征匹配的层次化语义-运动先验
Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM
- University of Georgia(佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
KYS-SLAM是ORB-SLAM3的扩展,用连续对应调制取代特征拒绝,通过层次化语义-运动先验计算对应代价,在固定配置下显著降低KITTI、EuRoC等数据集的轨迹误差。
AI中文摘要:
基于局部描述子的立体视觉SLAM系统面临语义模糊、实例级混淆和独立运动物体等问题,这些问题都会破坏数据关联并累积为轨迹漂移。现有的语义和动态SLAM方法通过二值特征拒绝来解决这些问题,以牺牲对应密度为代价进行离群点抑制。我们认为,上下文不可信性更适合表示为分级量而非排除标准。我们提出了Know-Your-Scene (KYS)-SLAM,这是ORB-SLAM3的一个模块化扩展,用连续对应调制取代特征拒绝。其贡献在于将上下文证据重新定义为对应代价,应用于特征匹配中,且不修改几何后端。每个关键点都通过层次化兼容性公式融合语义、全景和运动先验,其中语义类别和实例身份强制结构合理性,而零样本运动评分对独立运动物体上的特征进行降权。该评分来自一个免训练模块,该模块将深度感知的自我运动模型拟合到背景光流,并通过自校准、覆盖感知阈值对全景分割进行分类,因此只有具有足够运动证据的分割才会被惩罚,静态结构则不受惩罚。惩罚对应关系而非丢弃它们,保留了束调整所依赖的几何支持。在一种固定配置下,无需针对每个序列或数据集重新调整系数,KYS-SLAM在室外KITTI上将每序列ATE RMSE降低了17.4%,在室内EuRoC上降低了27.7%(共21个立体序列,无回归),在KITTI Tracking的动态子集上降低了6.6%,在Virtual KITTI 2上降低了17.8%(最高达31.2%)——在户外驾驶、室内飞行和合成图像等跨域场景中,仅用一组常数即可实现跨域迁移。
英文摘要:
Stereo visual SLAM systems built on local descriptors suffer from semantic ambiguity, instance-level confusion, and independently moving objects, each corrupting data association and accumulating as trajectory drift. Prevailing semantic and dynamic SLAM methods address this through binary feature rejection, sacrificing correspondence density for outlier suppression. We contend that contextual implausibility is better expressed as a graded quantity than an exclusion criterion. We present Know-Your-Scene (KYS)-SLAM, a modular extension of ORB-SLAM3 that supplants feature rejection with continuous correspondence modulation. The contribution is the reframing of contextual evidence as correspondence cost, applied within feature matching and leaving the geometric backend unmodified. Each keypoint is augmented with semantic, panoptic, and motion priors fused through a hierarchical compatibility formulation, in which semantic class and instance identity enforce structural plausibility while a zero-shot motion score down-weights features on independently moving objects. That score comes from a training-free module fitting a depth-aware ego-motion model to background optical flow and classifying panoptic segments via self-calibrating, coverage-aware thresholds, so only segments with sufficient motion evidence are penalized and static structure is left unpenalized. Penalizing correspondences rather than discarding them preserves the geometric support bundle adjustment depends on. Under one fixed configuration, no coefficient retuned per sequence or dataset, KYS-SLAM reduces per-sequence ATE RMSE by 17.4% on outdoor KITTI and 27.7% on indoor EuRoC across 21 stereo sequences with no regressions, and by 6.6% on dynamic subsets of KITTI Tracking and 17.8%, up to 31.2%, on Virtual KITTI 2 -- cross-domain transfer across outdoor driving, indoor flight, and synthetic imagery under one set of constants.