arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

看见不可见:物理引导的视觉提示用于温度和辐射感知的VLA导航

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

Hojoon Son, Fan Zhang

arXiv 2610.07558首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA导航中相机无法感知的温度和辐射风险,提出物理引导的视觉提示模块PG-VP,通过动态虚拟障碍物引导策略绕开不可见危险,无需重新训练,在仿真和真实实验中显著提升安全性。

AI 中文摘要

视觉-语言-动作(VLA)模型已成为视觉与语言导航(VLN)的主要范式。然而,在安全关键设施中,辐射或温度骤升等不可见风险无法被RGB相机检测到,而处理每种风险代价高昂,需要新的编码器、新数据以及模型重新训练。我们提出物理引导的视觉提示(PG-VP),这是一种即插即用的多模态感知模块,它复用冻结的VLA模型已擅长的能力:避开可见障碍物。给定近端辐射或热源,PG-VP执行物理引导的风险评估以确定避让方向,并叠加一个在连续帧间移动的相应虚拟障碍物(动态视觉提示)。导航策略随后自然绕开这一不可见危险。无论危险类型如何,均使用相同的虚拟障碍物,因此随着传感器的添加,视觉提示模式保持固定。当未检测到危险时,不渲染任何内容,策略行为与无PG-VP时完全一致。我们在OmniNav上使用R2R-CE和RxR-CE的val-unseen分割评估PG-VP,在84.9%和83.2%的情况下引导策略采取预期的低风险动作,代价是导航成功率分别下降6.8和7.9个百分点。我们还在真实机器人上针对不同场景,在存在实际热源和辐射源的情况下进行了测试,全程无需重新训练。真实测试表明,PG-VP有效避开这些不可见危险,针对热源和辐射源,最差10%平均轨迹安全性分别提升63.45%和32.59%。

英文摘要

Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and model retraining. We propose Physics-Guided Visual Prompting (PG-VP), a plug-and-play multimodal perception module that instead reuses what a frozen VLA model already does well: avoiding visible obstacles. Given a proximal radiation or thermal source, PG-VP performs a physics-guided risk assessment to determine the avoidance direction and overlays a corresponding virtual obstacle that moves across consecutive frames (Dynamic Visual Prompting). The navigation policy then naturally detours around this invisible hazard. The identical virtual obstacle is used regardless of hazard type, so the visual prompting pattern remains fixed as sensors are added. When no hazard is detected, nothing is rendered, and the policy behaves exactly as it would without PG-VP. We evaluate PG-VP on OmniNav using the val-unseen splits of R2R-CE and RxR-CE, where it guides the policy toward intended low-risk actions in 84.9% and 83.2% of cases, at a cost of 6.8 and 7.9 percentage points in navigation success rate. We further test it with distinct scenarios on a real robot in the presence of actual thermal and radiation sources, all without any retraining. The real test shows that PG-VP effectively avoids these invisible hazards, improving worst-10% average trajectory safety by 63.45% and 32.59% against thermal and radiation sources, respectively.

Comments8 pages, 7 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑