发表机构
Tianjin University; The University of Hong Kong; Hefei University of Technology(天津大学; 香港大学; 合肥工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对测试时自适应视觉语言导航中的分布偏移问题,提出信用引导策略改进(CGPI),从动作引起的观测转换中恢复决策信用以指导策略更新,并在冻结预训练策略下实现轻量级适应,在多个基准和机器人试验中取得一致提升。
AI 中文摘要
测试时自适应视觉语言导航(TTA-VLN)使预训练策略能够仅利用测试时的观测和交互历史在线适应未见环境。然而,分布偏移可能扭曲局部动作偏好并导致偏离路线的决策。现有方法依赖预测不确定性、轨迹级反馈或累积的适应经验来纠正此类偏差。然而,这些信号并不直接揭示所执行的动作是否支持指令引导的朝向目标的进展。此外,一个看似合理的纠正信号并不能保证可靠的策略更新。因此,关键挑战是双重的:识别支持目标导向改进的交互,并确定由此产生的更新是否值得保留。我们观察到,每个执行的动作都会引起即时的观测转换,提供其局部后果的证据。基于这一见解,我们提出了信用引导的策略改进(CGPI),它从动作引起的观测转换中恢复有符号的、参考相对的决策信用,而无需外部结果反馈。在冻结预训练导航策略的情况下,CGPI利用该信用提出轻量级的适应更新,并对照先前信用支持的交互进行验证。更新仅在得到支持时保留,否则回滚。CGPI在评估的VLN基准和导航骨干网络上取得了一致的性能提升,而定性机器人试验进一步展示了零样本模拟到现实迁移的可行性。
英文摘要
Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.