发表机构
Noah’s Ark Lab, 2012 Labs, Huawei(诺亚方舟实验室,2012实验室,华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言导航中标准行为克隆问题,提出FSC-VLN方法,通过添加未来查询令牌并经训练后移除的目标分支将其隐藏状态与未来视觉嵌入对齐,实验表明该方法在R2R val-unseen上提升了相关指标。
AI 中文摘要
端到端视觉语言导航(VLN)利用因果视觉语言模型可将指令和自我中心观察直接映射到行动,但标准行为克隆仅监督下一个行动,未明确训练策略状态以预测未来视觉结果。首先提出诊断性问题:若在训练和测试时给策略提供专家轨迹未来图像作为特权输入,该额外视觉证据对选择当前行动是否有用?答案是肯定的。接着提出可部署问题:推理时不访问未来图像,仅用压缩未来视觉潜在特征作为训练监督能否从未来信息中受益?提出了未来状态条件的VLN(FSC-VLN),在R2R val-unseen上,FSC-VLN在两种训练数据机制下比StreamVLN风格基线提高了SR/OSR/SPL,在长视野情节上增益更大;消融实验进一步支持了双查询设计。
英文摘要
End-to-end vision-language navigation (VLN) with causal vision-language models maps instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly encourage the policy state to be predictive of future visual outcomes, limiting long-horizon decision making. A privileged-input diagnostic shows that access to an expert-trajectory future image can substantially improve navigation, indicating that future observations contain rich, actionable cues, though such inputs are unavailable at deployment. Motivated by this signal, we propose Future-State-Conditioned VLN (FSC-VLN), a deployable model that augments a causal policy with a future-query token and uses training-only future-state supervision to distill information from future observations into the policy state. Concretely, during training we align the future-query representation to a frozen visual embedding $Δ$ steps ahead, while inference requires only past and current observations. This design preserves the baseline inference pattern and adds only two learned prefix tokens, implying minimal overhead. On R2R val-unseen, FSC-VLN improves SR/OSR/SPL over a StreamVLN-style baseline under two training-data regimes, with larger gains on long-horizon episodes; ablations further support the dual-query design that separates future and action queries.
Comments9 pages, 1 figure, 4 tables