arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪种地形更好?基于VLM原型的偏好学习用于越野可通行性排序

Which Terrain Is Better? Preference Learning with VLM Prototypes for Off-Road Traversability Ranking

Ji-Hoon Hwang, Jisung Bae, E-In Son, Dong-Wook Kim, Jung-Taak Kim, Seung-Woo Seo

arXiv 2609.23673首次发表:更新:

AI 中文总结

针对越野导航中地形偏好缺乏通用真值的问题,提出TravPro方法,利用VLM原型将稀疏比较转化为密集偏好伪标签,在五个未见领域达到0.915平均成对准确率,优于基线。

AI 中文摘要

在基于视觉的越野导航中,机器人不仅需要知道哪些障碍物需要避开,还需要知道哪种地形更好。前者通过自由空间检测或语义分割来处理。后者通常通过可通行性分数来回答,但这样的分数不存在通用的真值,因此感知系统退回到每个语义类别的预定义值或自由空间置信度。这些分数说明了某个区域是什么,而不是机器人应该偏好哪个区域。因此,我们将这种偏好表述为视觉可通行性排序,即对可见地形进行排序,可以通过两个区域之间的比较来进行监督。标准标注并不标记偏好,但它们隐含了偏好的方向。我们提出了TravPro,它将这些标注转换为有序的区域对,并在冻结的视觉-语言模型(VLM)补丁令牌上拟合一个小型读出器到这些区域对上。这些令牌被聚类一次到一个固定的原型库中,读出器为每个原型学习一个偏好分数。然后,读出器被应用于每个补丁,并充当教师,将稀疏的比较转换为密集的偏好伪标签,而无需像素级标注。一个RGB学生网络将这些映射蒸馏为密集的地形偏好图,同时生成一个非地面掩码,将障碍物和背景从排序中排除。在五个未见过的领域上,TravPro达到了0.915的平均成对准确率,而最强基线的准确率为0.783,产生了对表面条件敏感的排序,这是每类值无法表示的。当VLM被提示且监督被用作密集目标时,相同的VLM和相同的监督无法产生这样的排序;重要的是它们的使用方式。

英文摘要

In vision-based off-road navigation, a robot needs to know not only which obstacles to avoid but also which terrain is better. The first is handled by freespace detection or semantic segmentation. The second is usually answered with a traversability score, but no universal ground truth exists for such a score, so perception falls back on a predefined value per semantic class or a freespace confidence. These scores say what a region is, not which region a robot should prefer. We therefore formulate this preference as visual traversability ranking, an ordering of visible terrain that can be supervised by comparisons between two regions. Standard annotations do not label preference, but they imply its direction. We present TravPro, which converts these annotations into ordered region pairs and fits a small readout on frozen vision--language model (VLM) patch tokens to these pairs. The tokens are clustered once into a fixed prototype bank, and the readout learns a preference score per prototype. The readout is then applied to every patch and serves as a teacher that turns sparse comparisons into dense preference pseudo-labels without pixel-wise annotation. An RGB student distills these maps into a dense terrain-preference map together with a non-ground mask that excludes obstacles and background from the ranking. On five unseen domains, TravPro reaches a mean pairwise accuracy of 0.915 against 0.783 for the strongest baseline, producing an ordering sensitive to surface condition that a per-class value cannot represent. The same VLM and the same supervision yield no such ordering when the VLM is prompted and the supervision is used as dense targets; what matters is how they are used.

Comments8 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑