arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20983cs.RO

PIVOT:面向野外机器人导航的物理信息视觉-语言越野可通行性

PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation

Aoran Jiao, Wenda Zhao, Hshmat Sahak, Timothy D. Barfoot

首次发表
浏览论文内容

中文总结 AI 辅助

PIVOT系统融合几何规划与视觉-语言模型语义推理,通过物理相关性加权评分,在越野导航中将自主性从59.6%提升至97.0%,减少人工干预,显著增强自主导航能力。

中文摘要 AI 辅助

地形评估是越野移动机器人的关键能力,使其能够在非结构化和几何复杂的环境中实现安全可靠的导航。传统的基于几何的地形评估计算速度快,但在非结构化环境中往往过于保守。我们提出了PIVOT:一种物理信息视觉-语言越野可通行性导航系统,该系统将传统的基于几何的规划与基于视觉-语言模型(VLM)的语义推理相结合,用于野外机器人。为了在物理上夯实这一评估,我们量化了VLM预测的穿越能量成本、机器人振动和车轮打滑与真实世界测量的相关程度,并引入了一个统一的通行性评分,该评分根据每种模态的预测-测量相关性对其加权。为了提高效率,我们设计了一个两级导航架构,保留基于几何的规划作为标称模式,仅当该模式无法找到路径时才调用语义重新规划。在混合地形路线上进行的五次重复闭环试验中,总里程约6.4公里,与仅使用几何导航相比,所提出的系统将整体自主性从59.6%提高到97.0%,将人工干预次数从11次减少到3次,并将干预之间的平均距离(MDBI)从69.2米增加到412.9米。这些结果表明,基于物理的VLM地形评估可以大幅扩展自主导航超越仅几何的局限性,同时保留高效的几何规划作为标称模式。

英文摘要

Terrain assessment is a critical capability for off-road mobile robots, enabling safe and reliable navigation through unstructured and geometrically complex environments. Conventional geometry-based terrain assessment is fast to compute but often overly conservative in unstructured environments. We present PIVOT: a Physically Informed Vision-Language Off-Road Traversability navigation system that augments conventional geometry-based planning with vision-language-model (VLM)-based semantic reasoning for field robots. To physically ground this assessment, we quantify how strongly the VLM's predicted traversal energy cost, robot vibration, and wheel slip correlate with real-world measurements and introduce a unified traversability score that weights each modality by its prediction-measurement correlation. For efficiency, we design a two-level navigation architecture that retains geometry-based planning as the nominal mode and invokes semantic replanning only when that mode fails to find a path. Across five repeated closed-loop trials on a mixed-terrain route totalling around $6.4$ km, the proposed system increases overall autonomy from $59.6\%$ to $97.0\%$, reduces human interventions from $11$ to $3$, and increases the mean distance between interventions (MDBI) from $69.2$ m to $412.9$ m compared with geometry-only navigation. These results demonstrate that physically grounded VLM-based terrain assessment can substantially extend autonomous navigation beyond the limitations of geometry alone, while preserving efficient geometric planning as the nominal mode.

↑