AVERT-VLN:面向视觉与语言导航的弃权感知视觉错误恢复与训练
AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation
浏览论文内容
中文总结 AI 辅助
AVERT-VLN提出弃权感知的闭环框架,通过即插即用监控器在线请求人工恢复并离线偏好学习,在R2R-CE和RxR-CE上分别达到76.2%和66.3%的成功率。
中文摘要 AI 辅助
在未见环境中部署视觉与语言导航(VLN)智能体仍然具有挑战性,因为不熟悉的布局和视觉条件可能导致执行偏离轨道。与其依赖持续的人工监督,一种实用的策略是选择性地请求纠正性指导、恢复正在进行的任务,并重用纠正性交互以改进后续导航。我们提出了面向视觉与语言导航的弃权感知视觉错误恢复与训练(AVERT-VLN),这是一个闭环框架,使用即插即用的视觉语言监控器进行在线人工辅助恢复和离线偏好学习。该监控器独立于导航决策生成运行,并从指令、视觉历史和当前观察中评估指令-执行一致性。为了训练监控器识别偏差,我们构建了包含20K条反事实风险轨迹和基于规则的偏差标签的LOSTNAV数据集。监控器首先在40K条正常轨迹上进行微调以评估指令进度,然后在正常和风险轨迹上联合微调以识别语义偏差。在运行时,异步边车监控与导航模型并行评估执行情况。当控制器接受LOST判定时,它暂停自主执行并请求人工指导以进行恢复。对于离线策略改进,轨迹锚定偏好学习在共享决策上下文下将与偏差相关的失败转化为决策级偏好对,将监督限制在针对纠正的决策上。在人工辅助评估下,完整的AVERT-VLN系统在R2R-CE和RxR-CE的val-unseen分割上分别实现了76.2%和66.3%的成功率。相同的监控和人工辅助恢复接口也提高了三个评估导航架构的成功率。
英文摘要
Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.
发表机构
- Institute of Cyber-Systems and Control, Zhejiang University(浙江大学控制系统研究所)
机构由 AI 辅助整理,请以论文原文为准。