arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22678cs.ROcs.AIcs.CV

RACO:面向巡检的无人机视觉语言导航的可靠性感知粗目标优化

RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

Sen Wang, Yiming Sun, Jiaxuan He, Pengfei Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对面向巡检的无人机视觉语言导航,提出RACO框架,通过可靠性感知优化粗目标,在未见场景下较HETT基线显著提升成功率,改善巡检区域到达率并降低错误验证风险。

中文摘要 AI 辅助

无人机视觉语言导航(UAV-VLN)通常以到达目标作为评估标准,但面向巡检的部署要求智能体停留在有效的巡检区域内,并避免错误确认视觉或语义相似的干扰项。这一需求暴露了现有由粗到精的无人机视觉语言导航策略的一个关键弱点:在局部细化之前预测的粗目标常被视为可靠,尽管它可能漂移到看似合理但不正确的物体区域,限制了局部阶段的恢复能力。为了系统评估该问题,我们引入了LG-UVI,这是一种源自CityNav/CityRefer的以物体为中心的巡检评估设置。LG-UVI在标准无人机视觉语言导航任务中加入了目标物体、困难干扰项、类型感知巡检区域,以及巡检区域到达和物体级确认的诊断指标。为解决这一面向巡检的设置,我们进一步提出了RACO,一种可靠性感知的自适应由粗到精导航框架。RACO不将预测的粗目标视为固定航路点,而是将其视为运行时假设,在第1阶段之前及第1阶段到第2阶段的边界处,利用物体级候选锚点检查并修正粗定位。RACO还应用尺度自适应终端细化,通过运行时可观测的几何和基于锚点的证据处理终端近失情况。在统一的在线评估协议下,RACO在验证集未见和测试集未见场景下,较复现的HETT基线分别将成功率(SR)提升9.53和7.98个百分点。它还提高了巡检区域到达率,降低了错误验证风险,表明粗目标可靠性优化是对现有由粗到精无人机视觉语言导航策略的有效补充。

英文摘要

UAV vision-language navigation (UAV-VLN) is commonly evaluated as goal reaching, but inspection-oriented deployment requires the agent to stop within a valid inspection region and avoid falsely confirming visually or semantically similar distractors. This requirement exposes a key weakness in existing coarse-to-fine UAV-VLN policies: the coarse goal predicted before local refinement is often treated as reliable, although it may drift toward plausible but incorrect object regions and limit the ability of the local stage to recover. To systematically evaluate this problem, we introduce LG-UVI, an object-centric inspection evaluation setting derived from CityNav/CityRefer. LG-UVI extends standard UAV-VLN episodes with target objects, hard distractors, type-aware inspection regions, and diagnostics for inspection-region arrival and object-level confirmation. To address this inspection-oriented setting, we further propose RACO, a reliability-aware adaptive coarse-to-fine navigation framework. Instead of treating the predicted coarse goal as a fixed waypoint, RACO views it as a runtime hypothesis and uses object-level candidate anchors to check and correct coarse localization before Stage 1 and at the Stage 1-to-Stage 2 boundary. RACO also applies scale-adaptive terminal refinement to handle terminal near-miss cases using runtime-observable geometric and anchor-based evidence. Under a unified online evaluation protocol, RACO improves SR over the reproduced HETT baseline by 9.53 and 7.98 percentage points on validation-unseen and test-unseen, respectively. It also improves inspection-region arrival and reduces false verification risk, showing that coarse-goal reliability optimization is an effective complement to existing coarse-to-fine UAV-VLN policies.

↑